Back to News
bstnxbt (Open Source)May 3, 2026

dflash-mlx

Paid
Apple Silicon
MLX
Speculative Decoding
Local LLM
Inference
Qwen
Open Source

Explore dflash-mlx

Visit the official website to learn more and get started

### TL;DR

dflash-mlx brings lossless DFlash speculative decoding to Apple Silicon via MLX, achieving up to 4.37x token generation speedup while guaranteeing every emitted token is verified against the target model. Uses a small ~1B draft model to generate 16 tokens in parallel via block diffusion, then verifies them in a single forward pass. Ideal for local LLM inference on Mac hardware.

Key Insights & Metrics

Pricing
Free (Open Source)
Cost structure
Version
latest
Current release version
Hardware
Apple Silicon Mac (M1/M2/M3/M4/M5), macOS, MLX 0.31.1+
Compute requirements
Category
Paid
Licensing model
Region
Global
Primary region

Key Features

  • Up to 4.37x speedup on Qwen3.5-9B with 86-91% acceptance rates across all tested models
  • Lossless output: every token is verified via greedy acceptance before being committed — no hallucinated tokens
  • OpenAI-compatible server (dflash serve) supporting streaming, tool calls, and chat templates; works with OpenCode, aider, Continue, Open WebUI

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode