Back to News
Tongyi Lab (Alibaba Group)July 20, 2026

Qwen-Audio-3.0-TTS

Paid
Text-to-Speech
Multilingual
Voice Cloning
AI Research

Explore Qwen-Audio-3.0-TTS

Visit the official website to learn more and get started

### TL;DR

Qwen-Audio-3.0-TTS is a high-performance text-to-speech system from Alibaba's Tongyi Lab that supports 16 languages and various Chinese dialects. The model offers two primary variants, Flash and Plus, designed for real-time interaction and high-fidelity generation respectively, while featuring natural-language style control and robust voice cloning.

Key Insights & Metrics

Pricing
Plus: $0.20 per 10,000 characters & Flash: $15.00 per 1 million characters
Cost structure
Version
3.0
Current release version
Hardware
Cloud-based; No specific hardware required
Compute requirements
Category
Paid
Licensing model
Region
China
Primary region

Key Features

  • Support for 16 languages and 20 Chinese dialect regions
  • Dual-model architecture: Flash (low latency) and Plus (high-fidelity)
  • Natural-language style control for emotion, pace, and scenario
  • Fine-grained inline tag support for non-verbal cues like [giggles] or [gasp]
  • Robust voice cloning even from noisy or reverberant reference audio

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode