Advanced AI Text-to-Speech System Developed for Yoruba Language

Researchers have introduced TTSYoruba, a sophisticated rule-based concatenative diphone speech synthesizer specifically designed for the Yoruba language. This innovative system is already operational online as part of the YorubaName.com open dictionary, demonstrating its practical application in preserving and promoting African linguistic heritage. The development highlights a significant step forward in making digital tools more inclusive for low-resource languages.
The TTSYoruba system processes tone-marked Yoruba text inputs to generate audio output. Its core functionality relies on a meticulously hand-crafted phonological rule system applied to an extensive inventory of 651 diphone units. These units encompass five distinct tonal variants for every consonant-vowel combination in Yoruba, reflecting the language's complex tonal nature which is crucial for meaning.
The paper details the system's intricate phonological architecture, including its precise logic for tonal file selection and its solution to the challenging three-way nasal disambiguation problem (distinguishing oral /n/, nasalized vowels, and syllabic nasals). A notable orthographic contribution is the adoption of the caron and circumflex symbols as standard single-vowel contour tone markers, integrating them into the TTS normalization pipeline and the WriteYoruba keyboard input tool to standardize written representation of tones.
Performance evaluation involved a listener study with 50 participants, yielding detailed Mean Opinion Scores (MOS) that validate the system's effectiveness. This research not only advances text-to-speech technology but also provides a vital resource for the Yoruba-speaking population, spanning countries like Nigeria, Benin, and Togo, by enhancing digital accessibility and supporting language preservation efforts.
More in research
New AI Diagnostic Tool Aletheia Offers Offline Support for African Healthcare
Aletheia is an offline-first AI clinical decision support system specifically designed for low-resource healthcare settings across sub-Saharan Africa, addressing the critical lack…
New AfriSwitch Benchmark Reveals Major Gaps in AI Speech Recognition for African Code-Switched Languages
AfriSwitch is a new 61.36-hour benchmark dataset of human-transcribed, real-world code-switched speech across 16 African languages. It reveals that current AI speech recognition…
New TranslatePsy-AfriSLM Models Dramatically Improve African Language Translation for Low-Resource AI
TranslatePsy-AfriSLM introduces open-source machine translation resources for 19 Sub-Saharan African languages, including curated and synthetic data, and fine-tuned SLMs. These…
New AI Model Improves Poverty Mapping in Africa by Quantifying Uncertainty
A new machine learning method uses satellite imagery to predict poverty levels across Africa, providing crucial uncertainty estimates for policymakers. This innovation helps…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.