New Open-Source AI Models Boost Speech Recognition for African Languages

Researchers have introduced DONDO, a new collection of open-source automatic speech recognition (ASR) base models specifically designed for African languages. Built on the w2v-BERT 2.0 self-supervised speech encoder, DONDO includes twenty-one monolingual and five multilingual models, covering twenty-seven language varieties spoken across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe.
The models were primarily fine-tuned using read speech from religious texts. This approach was chosen due to the broad, license-clear, and orthographically consistent coverage these texts provide for languages that typically lack extensive transcribed audio datasets. The development involved a multi-step fine-tuning process, which first adapted a shared multilingual model at a higher learning rate before annealing it to achieve performance comparable to, and often surpassing, strong monolingual baselines.
A notable innovation is a lightweight language-conditioning mechanism. This allows a single multilingual model to be directed to a specific target language during inference by injecting a one-hot language identity as a prefix to the acoustic features. The annealed multilingual models achieved average word error rates (WER) between 10-13%, significantly closing the performance gap with monolingual models while consolidating multiple languages into a single checkpoint.
All DONDO models are released under the Apache-2.0 license on the Hugging Face KhayaAI organization. This permissive licensing allows for free use and fine-tuning, including for commercial applications. The languages covered by DONDO are estimated to be spoken by over one hundred million first-language speakers, with a substantially larger number including second-language users, highlighting the significant potential impact of these advancements for African communities.
More in research
New Study Questions Linguistic Relatedness as Key to Low-Resource African Language ASR
This research directly addresses the challenges of developing automatic speech recognition for low-resource African languages. The study utilized two Africa-centric datasets to…
South African PhD Research Develops Multilingual AI for Misinformation Detection in isiZulu and Sepedi
A South African PhD study has developed a multilingual AI framework to detect misinformation in isiZulu and Sepedi, addressing a critical gap where most AI tools only cover…
New AI Tool to Uncover Data Bias Offers Critical Safeguard for African Medical Deployments
A new tool designed to detect hidden biases in medical AI training data is particularly significant for African healthcare. With the continent's rapidly expanding medical AI…
Advanced AI Text-to-Speech System Developed for Yoruba Language
Researchers have developed TTSYoruba, an advanced text-to-speech system specifically for the Yoruba language, a major language spoken across West Africa. This innovation addresses…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.

