<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Tool Calling on 每日拍拍</title>
    <link>https://dailypypy.org/tags/tool-calling/</link>
    <description>Recent content in Tool Calling on 每日拍拍</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-tw</language>
    <copyright>© 2026 每日拍拍</copyright>
    <lastBuildDate>Thu, 20 Aug 2026 11:47:00 +0800</lastBuildDate><atom:link href="https://dailypypy.org/tags/tool-calling/index.xml" rel="self" type="application/rss+xml" />
    <follow_challenge>
      <feedId>155076163427069952</feedId>
      <userId>154825760438254592</userId>
    </follow_challenge>
    
    
    <item>
      <title>MLX-LM 本機 API Server：OpenAI 相容介面、Prompt Cache 與 Tool Calling</title>
      <link>https://dailypypy.org/learn/mlx-lm-local-api-server/</link>
      <pubDate>Thu, 20 Aug 2026 11:47:00 +0800</pubDate>
      
      <guid>https://dailypypy.org/learn/mlx-lm-local-api-server/</guid>
      <description>&lt;!---
1440x768
prompt: masterpiece, best quality, highres, clean anime illustration, japanese anime style, soft shading, flat color design, 1girl, black hair, green eyes, white off-shoulder shirt, black short skirt, wide shot, small figure, three-quarter side view, sitting cross-legged while typing on a slim laptop, focused curious half-smile, glancing toward viewer, pastel apricot background, subtle floating connected nodes and soft glowing rounded panels without symbols or text, neat composition, detailed eyes, cute and smart vibe, minimal background, polished illustration, no text
negative prompt: worst quality, bad eye, bad hand, extra limbs, manga, multiple views, monochrome, text, signature
dedup note: This is a service-integration follow-up to ../mlx-lm-local-models/. The older post covers model selection, CLI/Python generation, chat templates, streaming, and a local chat CLI. This article starts at the HTTP boundary: mlx_lm.server, OpenAI-compatible clients, health checks, automatic prefix/KV cache reuse, tool-calling compatibility, concurrency, observability, and safe network exposure. It does not repeat general inference, batch evaluation, embeddings, quantization, or LoRA workflows.
---&gt;
&lt;p&gt;





&lt;figure&gt;
    &lt;img class=&#34;my-0 rounded-md&#34; loading=&#34;lazy&#34; alt=&#34;featured&#34; src=&#34;./featured.png&#34; /&gt;

  
&lt;/figure&gt;
&lt;/p&gt;</description>
      <media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://dailypypy.org/learn/mlx-lm-local-api-server/featured.png" />
    </item>
    
  </channel>
</rss>
