AfricaDailyAI
← Back Home
ResearchEthiopiaEritreaPan-Africa96% confidence

New AI Model Significantly Boosts Performance for Ge'ez-Script African Languages

New AI Model Significantly Boosts Performance for Ge'ez-Script African Languages

Researchers have developed VEXMLM, a new variant of the XLM-R language model, specifically designed to improve performance for low-resource African languages that use the Ge'ez script, such as Amharic and Tigrinya. Existing multilingual models often struggle with these languages due to high rates of unknown words and inefficient subword fragmentation caused by tokenizers primarily optimized for Latin scripts.

VEXMLM addresses these challenges by incorporating language-specific SentencePiece tokenizers trained on carefully selected Amharic and Tigrinya corpora. Its vocabulary is expanded with 30,000 Ge'ez-script subwords, with their embeddings initialized through subword averaging. The model undergoes a two-stage training process, including masked language modeling and supervised fine-tuning for tasks like question answering, named entity recognition, and sentiment analysis.

The results show that VEXMLM significantly outperforms both the original XLM-R and Glot500 models across all evaluated tasks for Amharic and Tigrinya, with notable improvements in recognizing out-of-vocabulary entities. Critically, the enhancements achieved for Amharic and Tigrinya also positively impact performance for 17 other African languages, demonstrating a broader benefit for the continent's linguistic diversity.

This research highlights that adapting vocabularies and tokenizers offers a computationally efficient method to enhance multilingual models for underrepresented languages, rather than requiring complete retraining. This approach is crucial for fostering more inclusive and effective AI applications across Africa's diverse linguistic landscape, enabling better access to AI technologies for millions of speakers.

More in research

The dispatch

One email a day. The AI stories shaping Africa.

Rewritten for clarity, sourced always. No spam; unsubscribe anytime.