2026•bstnxbt (Open Source)
dflash-mlx vlatest
dflash-mlx brings lossless DFlash speculative decoding to Apple Silicon via MLX, achieving up to 4.37x token generation speedup while guaranteeing every emitted token is verified against the target model. Uses a small ~1B draft model to generate 16 tokens in parallel via block diffusion, then verifies them in a single forward pass. Ideal for local LLM inference on Mac hardware.
Up to 4.37x speedup on Qwen3.5-9B with 86-91% acceptance rates across all tested models
Lossless output: every token is verified via greedy acceptance before being committed — no hallucinated tokens
OpenAI-compatible server (dflash serve) supporting streaming, tool calls, and chat templates; works with OpenCode, aider, Continue, Open WebUI
PricingFree (Open Source)
Versionlatest