Back to News
ByteDance•June 23, 2026
Seed Audio 1.0
Paid
AI Audio
Text-to-Audio
Generative AI
Multimodal
### TL;DR
Seed Audio 1.0 is a unified, multimodal audio generation model designed to create complete audio scenes, including speech, music, ambient sound, and sound effects from a single prompt. It enables creators to direct complex audio environments with fine-grained timing control and voice consistency without the need for manual post-production stitching.
Key Insights & Metrics
Pricing
Pay-as-you-go: $0.15 per minute
Cost structure
Version
1.0
Current release version
Hardware
Cloud-based; No specific hardware required
Compute requirements
Category
Paid
Licensing model
Region
China
Primary region
Key Features
- Full-scene audio orchestration of dialogue, music, and ambient sound
- Zero-shot voice cloning from up to three short reference clips
- Multi-character dialogue generation with distinct emotional and stylistic delivery
- Prompt-level timing control with 100ms precision
- Long-form audio generation up to 2 minutes per pass with continuation support
- Cross-lingual synthesis supporting over 20 languages including English and Chinese
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!