Geometric Regularization Improves LLM Performance for African Languages

This research paper explores the performance disparities between large language models (LLMs) when processing low-resource languages compared to high-resource languages. The authors investigate these differences through the lens of representational geometry, analyzing how hidden representations within LLMs vary across 30 languages based on data availability. A key finding is that low-resource languages consistently exhibit 'representational degeneration' in the final layers of LLMs.
To address this issue, the study proposes and evaluates the use of geometric regularization techniques during continued pretraining (CPT). These regularization terms are designed to penalize and reduce the observed degeneration. The efficacy of this approach was tested by adapting nine base LLMs monolingually to ten different African languages.
The experiments demonstrated that geometric regularization successfully mitigates representational degeneration during CPT for these African languages. For larger models, a cosine similarity-based regularization technique showed marginal performance improvements over standard CPT, with more significant gains noted on particularly challenging tasks. This work establishes a measurable distinction in the representational geometry of low- and high-resource languages within LLMs.
Ultimately, the research validates targeted geometric intervention as a promising strategy for enhancing continued pretraining for low-resource languages. This has substantial implications for the development of more equitable and effective AI tools for diverse linguistic communities, especially in Africa where many languages fall into the low-resource category.
More in research
New AI Diagnostic Tool Aletheia Offers Offline Support for African Healthcare
Aletheia is an offline-first AI clinical decision support system specifically designed for low-resource healthcare settings across sub-Saharan Africa, addressing the critical lack…
New AfriSwitch Benchmark Reveals Major Gaps in AI Speech Recognition for African Code-Switched Languages
AfriSwitch is a new 61.36-hour benchmark dataset of human-transcribed, real-world code-switched speech across 16 African languages. It reveals that current AI speech recognition…
New TranslatePsy-AfriSLM Models Dramatically Improve African Language Translation for Low-Resource AI
TranslatePsy-AfriSLM introduces open-source machine translation resources for 19 Sub-Saharan African languages, including curated and synthetic data, and fine-tuned SLMs. These…
New AI Model Improves Poverty Mapping in Africa by Quantifying Uncertainty
A new machine learning method uses satellite imagery to predict poverty levels across Africa, providing crucial uncertainty estimates for policymakers. This innovation helps…
The dispatch
One email a day. The AI stories shaping Africa.
Rewritten for clarity, sourced always. No spam; unsubscribe anytime.