Back to News
bstnxbt (Open Source)•May 3, 2026
dflash-mlx
Paid
Apple Silicon
MLX
Speculative Decoding
Local LLM
Inference
Qwen
Open Source
### TL;DR
dflash-mlx brings lossless DFlash speculative decoding to Apple Silicon via MLX, achieving up to 4.37x token generation speedup while guaranteeing every emitted token is verified against the target model. Uses a small ~1B draft model to generate 16 tokens in parallel via block diffusion, then verifies them in a single forward pass. Ideal for local LLM inference on Mac hardware.
Key Insights & Metrics
Pricing
Free (Open Source)
Cost structure
Version
latest
Current release version
Hardware
Apple Silicon Mac (M1/M2/M3/M4/M5), macOS, MLX 0.31.1+
Compute requirements
Category
Paid
Licensing model
Region
Global
Primary region
Key Features
- Up to 4.37x speedup on Qwen3.5-9B with 86-91% acceptance rates across all tested models
- Lossless output: every token is verified via greedy acceptance before being committed — no hallucinated tokens
- OpenAI-compatible server (dflash serve) supporting streaming, tool calls, and chat templates; works with OpenCode, aider, Continue, Open WebUI
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!