Back to News
inclusionAI•November 24, 2025
Ming-omni-tts-16.8B-A3B
Open Source
Audio
### TL;DR
Ming-omni-tts-16.8B-A3B is a high-performance unified audio generation model developed by inclusionAI. It enables precise control over speech attributes and facilitates the synthesis of speech, environmental sounds, and music in a single channel.
Key Insights & Metrics
Pricing
The model is available under the Apache-2.0 license, indicating it is open-source and free to use.
Cost structure
Version
1.0
Current release version
Hardware
The model is available in Safetensors format, which is compatible with various hardware configurations. Specific hardware requirements are not specified.
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Fine-grained Vocal Control: Supports precise control over speech rate, pitch, volume, emotion, and dialect through simple commands.
- Intelligent Voice Design: Features 100+ premium built-in voices and supports zero-shot voice design via natural language descriptions.
- Immersive Unified Generation: Jointly generates speech, ambient sound, and music in a single channel, delivering a seamless auditory experience.
- High-efficiency Inference: Introduces a "Patch-by-Patch" compression strategy, reducing inference frame rate to 3.1Hz, enabling podcast-style audio generation.
- Professional Text Normalization: Accurately parses and narrates complex formats, including mathematical expressions and chemical equations.
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!