Back to News
Google•January 1, 2026
WaxalNLP
Open Source
Audio
### TL;DR
The WaxalNLP dataset is a comprehensive collection of audio data for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) tasks in 14 African languages. It aims to enhance the accuracy and fluency of speech and language technologies for underserved African languages and serves as a resource for digital preservation.
Key Insights & Metrics
Pricing
Free
Cost structure
Version
1.0.0
Current release version
Hardware
CPU only
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region
Key Features
- Contains approximately 1,250 hours of transcribed natural speech for ASR
- Includes about 240 hours of scripted natural speech for TTS
- Covers 14 African languages spoken by over 100 million people across 40 Sub-Saharan countries
- Acquired through partnerships with Makerere University, The University of Ghana, Digital Umuganda, and Media Trust
- Funded by Google and the Gates Foundation to be openly accessible
Discussion
1
Upvotes
0
Downvotes
1 review
Sign in to leave a review
Reviews
Upvote
asif@marktechpost.com • Jan 26, 2026
Quick vote