New Open-Source AI Models Boost Speech Recognition for African Languages
Researchers have introduced DONDO, a collection of open-source automatic speech recognition (ASR) base models specifically designed for African languages. Built on the w2v-BERT 2.0 self-supervised speech encoder, DONDO includes twenty-one monolingual and five multilingual models, covering twenty-seven language varieties across six African countries: Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe.
The models were primarily fine-tuned using read speech from religious texts, a strategic choice to overcome the common challenge of limited transcribed audio for many African languages. This approach ensures broad, license-clear, and orthographically consistent data. The fine-tuning process involved a two-step learning-rate annealing procedure, allowing the models to adapt effectively and, in some cases, outperform existing monolingual baselines.
A notable innovation is the lightweight language-conditioning mechanism. This feature enables a single multilingual checkpoint to be directed to a specific target language during inference by injecting a one-hot language identity. This significantly improves efficiency and broad applicability.
The annealed multilingual models achieved impressive average word error rates (WER) of 10-13%, effectively closing the performance gap with monolingual models while consolidating support for numerous languages into a single checkpoint. This development is crucial for advancing AI accessibility and utility across the continent.
All DONDO models are openly available on the Hugging Face KhayaAI organization under the Apache-2.0 license, allowing free use and further fine-tuning, including for commercial applications. These models are estimated to serve approximately one hundred million first-language speakers, with a much larger reach when second-language users are considered, paving the way for more inclusive AI-powered applications in Africa.
More in tools
New AI System Generates Natural-Sounding Yoruba Speech
A new rule-based AI speech synthesizer, TTSYoruba, has been developed for the Yoruba language, a significant step for African language technology. This system is deployed online…
M-PESA Ethiopia and Gebeya Launch AI Mini App for Mobile Users
M-PESA Ethiopia has partnered with local tech firm Gebeya to launch an AI Mini App, making various AI tools accessible to Ethiopian mobile users. This initiative aims to…
US and Morocco Launch Initiative for Advanced Drone Training and AI Innovation in North Africa
The United States and Morocco are collaborating to establish a drone academy and advanced military training center in Tan-Tan by 2030. This facility will specifically train…
Kenyan Startup Fikra API Democratizes AI Access for African Developers with Localized Payments and Pricing
Kenyan startup Fikra API has launched an AI inference platform specifically designed for African developers, addressing critical barriers like high costs, USD-only pricing, and…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.