Mercury 2
### TL;DR
Mercury 2 is the world's fastest reasoning LLM, built on Inception Labs' diffusion architecture. It generates tokens in parallel rather than sequentially, achieving over 1,000 tokens/sec on NVIDIA Blackwell GPUs — making reasoning-grade quality viable within real-time latency budgets for agents, voice, and search pipelines.
Key Insights & Metrics
Key Features
- 1,009 tokens/sec on NVIDIA Blackwell — parallel diffusion decoding delivers >5x faster generation than autoregressive models at the same quality tier
- Tunable reasoning with 128K context, native tool use, and schema-aligned JSON output — production-ready for agentic loops, coding, voice, and RAG pipelines
- OpenAI API-compatible — drop in as a replacement with no rewrites required; $0.25/M input · $0.75/M output
→ Related Releases
Protenix
Protenix is an open-source, trainable PyTorch implementation of AlphaFold 3, designed for high-accuracy biomolecular structure prediction. It aims to advance accessible and extensible research tools for the computational biology community. ([github.com](https://github.com/bytedance/Protenix?utm_source=openai))
Z-Image
Z-Image is an efficient image generation foundation model developed by Tongyi-MAI, designed to produce high-quality, diverse, and stylistically versatile images. It serves as a robust base for creators, researchers, and developers seeking advanced image generation capabilities.
Discussion
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!