<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>VLM on 每日拍拍</title>
    <link>https://dailypypy.org/tags/vlm/</link>
    <description>Recent content in VLM on 每日拍拍</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-tw</language>
    <copyright>© 2026 每日拍拍</copyright>
    <lastBuildDate>Thu, 10 Sep 2026 10:49:00 +0800</lastBuildDate><atom:link href="https://dailypypy.org/tags/vlm/index.xml" rel="self" type="application/rss+xml" />
    <follow_challenge>
      <feedId>155076163427069952</feedId>
      <userId>154825760438254592</userId>
    </follow_challenge>
    
    
    <item>
      <title>MLX-VLM 實戰：Apple Silicon 本地圖片問答與多模態模型</title>
      <link>https://dailypypy.org/learn/mlx-vlm-image-chat/</link>
      <pubDate>Thu, 10 Sep 2026 10:49:00 +0800</pubDate>
      
      <guid>https://dailypypy.org/learn/mlx-vlm-image-chat/</guid>
      <description>&lt;!---
1440x768
prompt: masterpiece, best quality, highres, clean anime illustration, japanese anime style, soft shading, flat color design, 1girl, black hair, green eyes, white off-shoulder shirt, black short skirt, slight dutch angle, upper body, one hand on chin thinking, curious gentle smile, looking at viewer, pastel buttercream yellow background, subtle floating translucent picture frames with abstract color blocks and soft lens shapes without symbols or text, neat composition, detailed eyes, cute and smart vibe, minimal background, polished illustration, no text
negative prompt: worst quality, bad eye, bad hand, extra limbs, manga, multiple views, monochrome, text, signature
dedup note: Planned 2026-09-10 topic `mlx-vlm-image-chat` passed a strict 157-post inventory scan. ../mlx-lm-local-models/ covers text-only local generation, ../mlx-lm-local-api-server/ covers serving text models, and ../mlx-lm-long-context-memory/ covers token and KV-cache tradeoffs. This follow-up instead uses the separate MLX-VLM package and focuses on image preprocessing, multimodal chat templates, single-image questions, multi-image comparison, grounding habits, and Vision Feature Cache. It does not repeat text-model installation, quantization, API serving, embeddings, RAG, or long-context tuning.
source check: Reviewed against the current official MLX-VLM README and usage guide plus the selected Hugging Face model metadata on 2026-09-10. PyPI reported mlx-vlm 0.7.0 with Python &gt;=3.10; the tutorial keeps the model ID configurable because supported checkpoints and APIs continue to evolve.
---&gt;
&lt;p&gt;





&lt;figure&gt;
    &lt;img class=&#34;my-0 rounded-md&#34; loading=&#34;lazy&#34; alt=&#34;featured&#34; src=&#34;./featured.png&#34; /&gt;

  
&lt;/figure&gt;
&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://dailypypy.org/learn/mlx-vlm-image-chat/featured.png" />
    </item>
    
  </channel>
</rss>
