New Language Model VEXMLM Boosts AI Performance for Ethiopian and Eritrean Languages

Researchers have developed VEXMLM, a new variant of the multilingual language model XLM-R, specifically designed to improve natural language processing for low-resource Ge'ez-script languages like Amharic (Ethiopia) and Tigrinya (Ethiopia and Eritrea). Traditional large language models often struggle with these languages because their Latin-script-centric tokenizers inefficiently break down Ge'ez words into many subwords, hindering performance.
VEXMLM addresses this by incorporating language-specific SentencePiece tokenizers and expanding XLM-R's vocabulary with 30,000 Ge'ez-script subwords. This vocabulary expansion significantly reduces tokenizer 'fertility' (the number of subword tokens per word), making the processing of Amharic and Tigrinya text much more efficient. The model then undergoes a two-stage training process, including continued masked language modeling and fine-tuning on tasks like question answering, named entity recognition, and sentiment analysis.
While VEXMLM shows modest improvements in named entity recognition and comparable performance in sentiment analysis, its primary breakthrough lies in the efficiency of Ge'ez-script tokenization. The study highlights that while vocabulary expansion is key for efficiency, continued pretraining is essential for the expanded model to surpass baseline performance on downstream tasks. This research provides valuable open-source resources, including a GitHub repository and Hugging Face model, fostering further development in African language AI.
More in research
Ethiopian Researchers Develop AI for Early Breast Cancer Detection in Low-Resource Settings
Ethiopian researchers have developed an AI model, HCMAN, specifically designed for early breast cancer detection in low-resource settings across Sub-Saharan Africa. The model…
New MGhana-ST Dataset Boosts Speech Translation for Ghanaian Languages
This research introduces MGhana-ST, a new speech translation dataset for four low-resource Ghanaian languages: Ga, Twi, Ewe, and Fante. The dataset and accompanying analysis are…
AI Models Show Promise for Tuberculosis Screening in Uganda and South Africa
This research evaluates machine learning models for tuberculosis screening using clinical data gathered in Uganda and South Africa. The study demonstrates the viability of…
New AI System Delivers High-Quality Yoruba Speech Synthesis
A new rule-based AI speech synthesizer, TTSYoruba, has been developed specifically for the Yoruba language, addressing its complex tonal and phonetic features. This system is…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.