Back to News
Tongyi Lab (Alibaba Group)•July 20, 2026
Qwen-Audio-3.0-TTS
Paid
Text-to-Speech
Multilingual
Voice Cloning
AI Research
### TL;DR
Qwen-Audio-3.0-TTS is a high-performance text-to-speech system from Alibaba's Tongyi Lab that supports 16 languages and various Chinese dialects. The model offers two primary variants, Flash and Plus, designed for real-time interaction and high-fidelity generation respectively, while featuring natural-language style control and robust voice cloning.
Key Insights & Metrics
Pricing
Plus: $0.20 per 10,000 characters & Flash: $15.00 per 1 million characters
Cost structure
Version
3.0
Current release version
Hardware
Cloud-based; No specific hardware required
Compute requirements
Category
Paid
Licensing model
Region
China
Primary region
Key Features
- Support for 16 languages and 20 Chinese dialect regions
- Dual-model architecture: Flash (low latency) and Plus (high-fidelity)
- Natural-language style control for emotion, pace, and scenario
- Fine-grained inline tag support for non-verbal cues like [giggles] or [gasp]
- Robust voice cloning even from noisy or reverberant reference audio
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!