New Open-Source AI Models Boost Speech Recognition for African Languages

Researchers have introduced DONDO, a collection of open-source automatic speech recognition (ASR) base models specifically designed for African languages. Built on the w2v-BERT 2.0 self-supervised speech encoder, DONDO includes twenty-one monolingual and five multilingual models, covering twenty-seven language varieties across six African countries: Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe.
The models were primarily fine-tuned using read speech from religious texts, a strategic choice to overcome the common challenge of limited transcribed audio for many African languages. This approach ensures broad, license-clear, and orthographically consistent data. The fine-tuning process involved a two-step learning-rate annealing procedure, allowing the models to adapt effectively and, in some cases, outperform existing monolingual baselines.
A notable innovation is the lightweight language-conditioning mechanism. This feature enables a single multilingual checkpoint to be directed to a specific target language during inference by injecting a one-hot language identity. This significantly improves efficiency and broad applicability.
The annealed multilingual models achieved impressive average word error rates (WER) of 10-13%, effectively closing the performance gap with monolingual models while consolidating support for numerous languages into a single checkpoint. This development is crucial for advancing AI accessibility and utility across the continent.
All DONDO models are openly available on the Hugging Face KhayaAI organization under the Apache-2.0 license, allowing free use and further fine-tuning, including for commercial applications. These models are estimated to serve approximately one hundred million first-language speakers, with a much larger reach when second-language users are considered, paving the way for more inclusive AI-powered applications in Africa.
More in tools
Viamo Launches Offline AI Voice Platform in Ghana to Bridge Digital Divide
Viamo has launched the 231 Voice Platform in Ghana, offering offline AI access via toll-free phone calls to millions without smartphones or internet. This initiative uses basic…
Tether AI Unveils Offline Translation Models for 19 African Languages, Benchmarked by Peer Review
Tether AI Research has released open-source, offline translation models for 19 African languages, designed to run on basic smartphones without internet access. This addresses…
New Offline AI Diagnostic Tool Promises to Transform Healthcare in Rural Sub-Saharan Africa
Aletheia is an offline-first AI clinical decision support system specifically designed for low-resource healthcare settings in sub-Saharan Africa. It was fine-tuned on a dataset…
Intron's Sahara v2.5 Enhances African Language AI with Mid-Sentence Code-Switching and Igbo/Hausa Voice Generation
Nigerian voice AI company Intron has released Sahara v2.5, which now supports mid-sentence language switching across around 20 African languages and offers voice generation in…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.


