<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Benchmark on 每日拍拍</title>
    <link>https://dailypypy.org/tags/benchmark/</link>
    <description>Recent content in Benchmark on 每日拍拍</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-tw</language>
    <copyright>© 2026 每日拍拍</copyright>
    <lastBuildDate>Mon, 28 Sep 2026 10:44:00 +0800</lastBuildDate><atom:link href="https://dailypypy.org/tags/benchmark/index.xml" rel="self" type="application/rss+xml" />
    <follow_challenge>
      <feedId>155076163427069952</feedId>
      <userId>154825760438254592</userId>
    </follow_challenge>
    
    
    <item>
      <title>MLX-LM Speculative Decoding：Draft Model、加速條件與品質 Benchmark</title>
      <link>https://dailypypy.org/learn/mlx-lm-speculative-decoding/</link>
      <pubDate>Mon, 28 Sep 2026 10:44:00 +0800</pubDate>
      
      <guid>https://dailypypy.org/learn/mlx-lm-speculative-decoding/</guid>
      <description>&lt;!---
1440x768
prompt: masterpiece, best quality, highres, clean anime illustration, japanese anime style, soft shading, flat color design, 1girl, black hair, green eyes, white off-shoulder shirt, black short skirt, wide shot, small figure, three-quarter view, standing with hands behind back between one small translucent glowing orb and one larger translucent glowing orb, bright analytical smile, looking at viewer, pastel rose quartz and soft teal background, subtle paired flowing light trails converging ahead without symbols, logos, markings, numbers, letters, or text, neat composition, detailed eyes, cute and smart vibe, minimal background, polished illustration, no text
negative prompt: worst quality, bad eye, bad hand, extra limbs, manga, multiple views, monochrome, text, signature
dedup note: This is a focused follow-up to ../mlx-lm-local-models/, ../mlx-lm-batch-inference/, ../mlx-lm-quantization-convert/, and ../mlx-lm-long-context-memory/. Those posts cover basic generation, repeatable prompt evaluation, weight quantization, and context/cache memory. This article instead centers target/draft compatibility, speculative token acceptance, draft-length tuning, end-to-end latency and throughput measurement, break-even analysis, and greedy quality parity. It does not repeat installation tours, model conversion recipes, generic sampling advice, or long-context cache tuning.
source check: Reviewed against the current official MLX-LM generate.py, server documentation, release history, and public issue tracker on 2026-09-28. The current generate CLI exposes --draft-model and --num-draft-tokens; stream_generate requires the same tokenizer, and speculative decoding requires a trimmable prompt cache. Public reports also show that model/cache combinations and greedy-output parity must be verified on the exact installed version rather than assumed.
---&gt;
&lt;p&gt;





&lt;figure&gt;
    &lt;img class=&#34;my-0 rounded-md&#34; loading=&#34;lazy&#34; alt=&#34;featured&#34; src=&#34;./featured.png&#34; /&gt;

  
&lt;/figure&gt;
&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://dailypypy.org/learn/mlx-lm-speculative-decoding/featured.png" />
    </item>
    
  </channel>
</rss>
