<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>OCR on 每日拍拍</title>
    <link>https://dailypypy.org/tags/ocr/</link>
    <description>Recent content in OCR on 每日拍拍</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-tw</language>
    <copyright>© 2026 每日拍拍</copyright>
    <lastBuildDate>Mon, 05 Oct 2026 10:31:00 +0800</lastBuildDate><atom:link href="https://dailypypy.org/tags/ocr/index.xml" rel="self" type="application/rss+xml" />
    <follow_challenge>
      <feedId>155076163427069952</feedId>
      <userId>154825760438254592</userId>
    </follow_challenge>
    
    
    <item>
      <title>MLX-VLM PDF 文件擷取：逐頁影像、結構化欄位與引用抽查</title>
      <link>https://dailypypy.org/learn/mlx-vlm-pdf-document-extraction/</link>
      <pubDate>Mon, 05 Oct 2026 10:31:00 +0800</pubDate>
      
      <guid>https://dailypypy.org/learn/mlx-vlm-pdf-document-extraction/</guid>
      <description>&lt;!---
1440x768
prompt: masterpiece, best quality, highres, clean anime illustration, japanese anime style, soft shading, flat color design, 1girl, black hair, green eyes, white off-shoulder shirt, black short skirt, wide shot, seated sideways on a low stool, carefully holding one translucent blank page panel with both hands, attentive gentle expression, looking at viewer, pastel dusty rose and pale cyan background, subtle layered blank paper sheets and small glowing anchor dots without symbols, logos, markings, numbers, letters, or text, neat composition, detailed eyes, cute and smart vibe, minimal background, polished illustration, no text
negative prompt: worst quality, bad eye, bad hand, extra limbs, manga, multiple views, monochrome, text, signature
dedup note: Planned 2026-10-05 topic `mlx-vlm-pdf-document-extraction` passed a strict 181-post inventory and current-source scan. ../mlx-vlm-image-chat/ covers interactive image questions, ../mlx-vlm-batch-structured-output/ covers independent image manifests and generic schema validation, and ../mlx-vlm-video-frame-analysis/ covers timestamped video frames. This focused follow-up instead centers PDF page rasterization, page anchors, native-text and OCR baselines, reading order, page-level evidence citations, cross-page conflict handling, and risk-based human sampling. It does not repeat the older image-chat, generic batch-runner, or video-timeline tutorials.
source check: Reviewed against the current official MLX-VLM README, usage guide, and v0.7.4 release information, plus the PyMuPDF 1.28.2 image, text, OCR, and Page documentation on 2026-10-05. MLX-VLM currently documents multimodal JSON Schema output through its OpenAI-compatible server. PyMuPDF documents DPI-based page rasterization, position-aware text blocks/words, and selective OCR through a reusable TextPage.
---&gt;
&lt;p&gt;





&lt;figure&gt;
    &lt;img class=&#34;my-0 rounded-md&#34; loading=&#34;lazy&#34; alt=&#34;featured&#34; src=&#34;./featured.png&#34; /&gt;

  
&lt;/figure&gt;
&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://dailypypy.org/learn/mlx-vlm-pdf-document-extraction/featured.png" />
    </item>
    
  </channel>
</rss>
