Back to News
inclusionAI•August 9, 2026
LLaDA2.2-flash
Open Source
diffusion
llm
text-generation
moe
### TL;DR
LLaDA2.2-flash is an agent-oriented Mixture-of-Experts (MoE) diffusion language model designed for long-context agentic workloads. It introduces Levenshtein Editing with DELETE and INSERT control tokens to enable efficient parallel generation, error correction, and multi-turn tool use.
Key Insights & Metrics
Pricing
Open source under Apache License 2.0
Cost structure
Version
2.2
Current release version
Hardware
High-performance GPU infrastructure recommended for 100B parameter MoE model; supports SGLang serving for optimized inference
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- 128K context window with Block Routing for efficient MoE expert activation
- Levenshtein Editing using DELETE and INSERT control tokens for structural sequence modification
- Agentic Reinforcement Learning via L-EBPO for improved error correction
- 100B parameter Mixture-of-Experts (MoE) architecture
- High-throughput parallel diffusion decoding
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!