New Open-Source AI Models Boost Speech Recognition for African Languages

Researchers have introduced DONDO, a new collection of open-source automatic speech recognition (ASR) base models specifically designed for African languages. Built on the w2v-BERT 2.0 self-supervised speech encoder, DONDO includes twenty-one monolingual and five multilingual models, covering twenty-seven language varieties spoken across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe.
The models were primarily fine-tuned using read speech from religious texts. This approach was chosen due to the broad, license-clear, and orthographically consistent coverage these texts provide for languages that typically lack extensive transcribed audio datasets. The development involved a multi-step fine-tuning process, which first adapted a shared multilingual model at a higher learning rate before annealing it to achieve performance comparable to, and often surpassing, strong monolingual baselines.
A notable innovation is a lightweight language-conditioning mechanism. This allows a single multilingual model to be directed to a specific target language during inference by injecting a one-hot language identity as a prefix to the acoustic features. The annealed multilingual models achieved average word error rates (WER) between 10-13%, significantly closing the performance gap with monolingual models while consolidating multiple languages into a single checkpoint.
All DONDO models are released under the Apache-2.0 license on the Hugging Face KhayaAI organization. This permissive licensing allows for free use and fine-tuning, including for commercial applications. The languages covered by DONDO are estimated to be spoken by over one hundred million first-language speakers, with a substantially larger number including second-language users, highlighting the significant potential impact of these advancements for African communities.
More in research
New AI Diagnostic Tool Aletheia Offers Offline Support for African Healthcare
Aletheia is an offline-first AI clinical decision support system specifically designed for low-resource healthcare settings across sub-Saharan Africa, addressing the critical lack…
New AfriSwitch Benchmark Reveals Major Gaps in AI Speech Recognition for African Code-Switched Languages
AfriSwitch is a new 61.36-hour benchmark dataset of human-transcribed, real-world code-switched speech across 16 African languages. It reveals that current AI speech recognition…
New TranslatePsy-AfriSLM Models Dramatically Improve African Language Translation for Low-Resource AI
TranslatePsy-AfriSLM introduces open-source machine translation resources for 19 Sub-Saharan African languages, including curated and synthetic data, and fine-tuned SLMs. These…
New AI Model Improves Poverty Mapping in Africa by Quantifying Uncertainty
A new machine learning method uses satellite imagery to predict poverty levels across Africa, providing crucial uncertainty estimates for policymakers. This innovation helps…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.