Back to News
AIDC-AI•May 1, 2026
LongSpeech
Paid
audio-LLM
dataset
speech
ASR
benchmarking
ICASSP2026
open-source
### TL;DR
LongSpeech is a large-scale long-form audio understanding dataset with 100,000+ audio segments (~10 min each), designed to benchmark and train Audio LLMs on long-form speech. Presented at ICASSP 2026, it covers 8 tasks: ASR, translation, summarization, speaker counting, QA, and emotion analysis.
Key Insights & Metrics
Pricing
Free
Cost structure
Version
1.0
Current release version
Hardware
Standard audio processing hardware
Compute requirements
Category
Paid
Licensing model
Region
China
Primary region
Key Features
- 100,000+ long-form audio segments (~10 minutes each) for benchmarking Audio LLMs
- Covers 8 tasks: ASR, translation, summarization, speaker counting, QA, emotion analysis
- Publicly released on Hugging Face with accompanying ICASSP 2026 paper
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!