Back to News
Microsoft•June 2, 2026
MAI-Voice-2
Paid
text-to-speech
AI model
Microsoft AI
voice synthesis
### TL;DR
MAI-Voice-2 is Microsoft's advanced text-to-speech model that converts text into expressive, natural-sounding speech across 15 languages. It offers realistic expression, instant voice matching, and is designed for long-form content like audiobooks and podcasts.
Key Insights & Metrics
Pricing
$22 per million characters
Cost structure
Version
2
Current release version
Hardware
A fully cloud hosted model through Microsoft Foundry with no local hardware requirements
Compute requirements
Category
Paid
Licensing model
Region
United States
Primary region
Key Features
- Realistic expression with organic pacing, tone, and emotional range
- Instant voice matching from a short reference clip
- Stable, high-fidelity output for long-form content
- Support for 15 languages with code-switching capabilities
- Granular emotion control via emotion tags
- Consent guardrails ensuring only authorized voices are used
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!