Uhura
A benchmark for testing scientific reasoning and truthfulness across six African languages.
Why Uhura
Language models are commonly evaluated in English and a small group of other high-resource languages. That leaves major blind spots in how reliably these systems reason, answer technical questions, and avoid false claims for much of the world.
Uhura makes that gap measurable. It provides a shared benchmark for asking whether a model remains useful and truthful when the language changes, not only when the question does.
Two tests, six languages
The benchmark pairs two complementary evaluations. Uhura-ARC-Easy measures multiple-choice scientific question answering. Uhura-TruthfulQA tests whether models repeat common misconceptions across sensitive subjects including health, law, finance, and politics.
Each task was translated by people into Amharic, Hausa, Northern Sotho, Swahili, Yoruba, and Zulu, creating comparable evaluations across typologically diverse languages.
- Scientific reasoning through Uhura-ARC-Easy
- Factual reliability through Uhura-TruthfulQA
- Human translation across six African languages

Translation is part of the research
Technical benchmarks cannot be translated safely by replacing words one at a time. Scientific terminology, regional usage, answer choices, and culturally specific assumptions all need review in context.
Uhura documents these difficulties and the mitigation work needed to make multilingual evaluation meaningful. The result is both a benchmark and a practical account of how to build better low-resource language datasets.
What the results reveal
Across the evaluated systems, performance was consistently stronger in English than in the African languages. The study also found a significant gap between leading proprietary systems and the open models included in the evaluation.
The central finding is straightforward: strong English performance does not guarantee safe, reliable behavior elsewhere. Multilingual capability needs to be tested directly and improved continuously before models are trusted in real-world settings.
Built to be extended
The paper, benchmark datasets, and evaluation platform were released openly to support further work in low-resource language NLP. Researchers can evaluate new models, examine individual languages, and build broader multilingual benchmarks from the same foundation.
Uhura is a starting point rather than a finish line: a way to make language equity visible, testable, and actionable.