Back to News
Inception LabsMay 12, 2026

Mercury 2

Paid
LLM
Diffusion Model
Fast Inference
Agentic
API
Reasoning
Real-time AI

Explore Mercury 2

Visit the official website to learn more and get started

### TL;DR

Mercury 2 is the world's fastest reasoning LLM, built on Inception Labs' diffusion architecture. It generates tokens in parallel rather than sequentially, achieving over 1,000 tokens/sec on NVIDIA Blackwell GPUs — making reasoning-grade quality viable within real-time latency budgets for agents, voice, and search pipelines.

Key Insights & Metrics

Pricing
$0.25/M input · $0.75/M output tokens
Cost structure
Version
Mercury 2
Current release version
Hardware
Cloud API — no local hardware required
Compute requirements
Category
Paid
Licensing model
Region
Global
Primary region

Key Features

  • 1,009 tokens/sec on NVIDIA Blackwell — parallel diffusion decoding delivers >5x faster generation than autoregressive models at the same quality tier
  • Tunable reasoning with 128K context, native tool use, and schema-aligned JSON output — production-ready for agentic loops, coding, voice, and RAG pipelines
  • OpenAI API-compatible — drop in as a replacement with no rewrites required; $0.25/M input · $0.75/M output

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode