AI That Speaks Our Languages: The Fight for African NLP
Ask a large language model a question in fluent Yoruba and watch what happens. Often it answers in mangled grammar, invents words, or quietly switches to English. Now consider that Yoruba has more than 40 million speakers. This is not a rounding error. It is the central injustice of the current AI wave: the technology reshaping the global economy barely speaks the languages of a billion Africans. Fixing that is not charity. It is the single highest-leverage AI project on the continent, and Africans are the only ones who will do it properly.
What "low-resource" really means
In machine-learning jargon, most African languages are "low-resource." The phrase sounds technical and neutral. It is neither. It means there is little digitised text to train on - few books scanned, few websites, few labelled datasets - not because the languages are small, but because the internet was built somewhere else. Swahili has over 100 million speakers across East Africa. Amharic, Hausa, isiZulu and Igbo each count tens of millions. Yet all of them are starved of the training data that English and Mandarin drown in.
The consequence compounds. A model that cannot read a language cannot serve its speakers, so those speakers generate less digital text in it, so the next model is even worse. Left alone, AI does not just ignore African languages. It actively accelerates their disappearance from digital life.
Why the global giants keep getting it wrong
The big labs are aware of the gap and have thrown scale at it. Google released a model in 2022 claiming support for more than 1,000 languages, adding several African ones to Google Translate. The problem is that scale alone does not buy quality. Google's own researchers admitted that native speakers rated many of the low-resource translations very poorly.
Here is the uncomfortable truth the benchmark charts hide: a small, focused team often beats a giant. The Ethiopian startup Lesan built machine translation for Amharic and Tigrinya that outperformed both Google Translate and Microsoft Translator across every language pair tested. Their edge was not a bigger GPU cluster. It was fieldwork - digitising offline print resources, building benchmarks from real Ethiopian text, and having people who actually speak the languages judge the output. Lesan has served tens of thousands of users and millions of translations. That is what domain ownership looks like.
Masakhane: research by Africans, for Africans
The beating heart of this movement is Masakhane - isiZulu for "we build together." It is a grassroots research collective with more than 2,000 members, most of them African, doing what they call participatory research. Native speakers, dataset curators, linguists and engineers work side by side, because you cannot label data in a language you do not speak, and you cannot judge a translation you cannot read.
Masakhane has produced benchmark datasets and models for dozens of African languages, and it feeds a wider ecosystem: the AfricaNLP workshop series, the pan-African Deep Learning Indaba, and Data Science Africa. This is not a single product. It is an entire research culture being built from scratch, decentralised across the continent, in defiance of a field that assumed the work was too niche to matter.
Lelapa AI and the dung-beetle strategy
In September 2024 the Johannesburg company Lelapa AI released InkubaLM, billed as a multilingual language model built for African languages - starting with Swahili, Yoruba, isiXhosa, Hausa and isiZulu. The most important detail is its size: about 0.4 billion parameters, tiny by frontier standards, yet competitive on its target languages.
It is named after the dung beetle, and the metaphor is the whole strategy. A dung beetle is small and does an enormous amount of work relative to its size. A 400-million-parameter model can run cheaply, even on modest hardware, close to where users actually are - which matters enormously in markets where GPU time is scarce and expensive.
The African NLP bet is not to out-scale OpenAI. It is to out-focus it - smaller models, better data, native judgement, and deployment that survives a weak connection and a tight budget.
Voice is the real frontier
There is a second reason African-language AI matters more than the text-translation debate suggests, and it is about literacy and interface. A large share of the continent's population is more comfortable speaking than reading, and many are more fluent in a local language than in the official one printed on government forms. A text-only, English-only product quietly excludes them twice over.
Voice changes that equation. A farmer who cannot read can still ask a question out loud and hear an answer back. A trader can check a balance by speaking. This is exactly why small, efficient models that run close to the device matter so much: voice interfaces need to respond fast and work on cheap hardware, and a bloated frontier model piped in from another continent cannot reliably do either. The teams building African speech datasets - often through the same Masakhane and AfricaNLP communities, and increasingly funded through efforts like the African Language Hub for AI - are laying track for products that have not been built yet. Whoever owns good local speech data owns the interface layer for the next hundred million users.
Why this is an economic argument, not a cultural one
It is tempting to frame African-language AI as a heritage project - important, but soft. That framing undersells it. Consider what language actually gates:
- Health. A patient who cannot describe symptoms to a triage system in their own language gets worse care. Voice interfaces in local languages reach people who cannot read English at all.
- Finance. Voice-first banking in Hausa or Swahili reaches customers that text-and-English interfaces lock out entirely.
- Agriculture. Nuru already delivers crop-disease advice in Swahili. The advice is useless if the farmer cannot understand it.
- Government and education. Public services delivered only in colonial languages exclude the majority by design.
Every one of these is a market. Language is not the decoration on top of the product - it is the access gate to hundreds of millions of customers that English-only competitors cannot reach.
What a developer can actually do
This is one of the rare AI frontiers where an individual contribution still moves the needle. Concretely: contribute to Masakhane's datasets in a language you speak. Test whether the model you are about to ship actually works in your users' language before you assume it does. If you are building voice or chat for an African market, evaluate InkubaLM and other local models against the default of a foreign API - the cheaper, smaller option may also be the better one, and it is exactly the kind of decision a business should make on evidence rather than hype, the same discipline that stops teams from buying AI before your business is ready. And record and donate speech and text data ethically, with consent, so the next model is less impoverished than the last.
The languages of Africa were left out of the last technology wave by accident and neglect. Whether they are left out of this one is now a choice - and, for the first time, it is largely an African choice to make.