
A Namibian computer science researcher is part of a University of Cape Town team that has developed a new artificial intelligence language model trained specifically on South Africa’s 11 official written languages, in a move aimed at addressing the exclusion of African languages from mainstream AI systems.
The research, led by Namibian master’s student Anri Lombard together with Jan Buys and Francois Meyer, will be presented at the Language Resources and Evaluation Conference 2026 in Mallorca this month.
The project introduces two systems: MzansiText, a multilingual dataset covering South Africa’s 11 official written languages, and MzansiLM, a language model trained from scratch using the dataset.
The development comes as AI tools such as chatbots and digital assistants continue to expand globally, while many African languages remain poorly supported due to limited training data.
According to the researchers, nine of South Africa’s 11 official written languages are still classified as low-resource languages because of the shortage of large digital text datasets required to train AI systems.
“In language modelling, languages are considered low resource, primarily because there are much fewer and smaller textual datasets available in these languages for training language models,” Buys said.
“Our dataset, MzansiText, is still small compared to data available for high-resource languages such as English and major European and Asian languages, but larger than previous datasets for South African languages.”
The researchers said MzansiLM is believed to be the first publicly available decoder-only language model designed specifically to support all 11 South African written languages.
“There has been real progress in language modelling for African languages, including some South African ones like isiXhosa and isiZulu,” Meyer said.
“But most existing models only cover a subset of languages. With MzansiLM, we wanted to build a single model focused specifically on South Africa that covers all 11 official written languages, including those that are often left out.”
Lombard said the project emerged from his master’s research into AI systems for low-resource languages, where he identified gaps in publicly available African language models.
“I came into this work through my master’s research, which looks at how different language-model architectures perform for low-resource languages, since that is still a relatively underexplored area,” Lombard said.
“One thing that stood out to me is that publicly available models tended to cover only a subset of the South African languages we care about. MzansiLM was meant to provide a small decoder-only baseline that future work can compare against and build on.”
Although the model contains 125 million parameters, far smaller than commercial AI systems such as ChatGPT, the team said tests showed it performed competitively in several African language tasks.
The researchers said the model outperformed larger open-source systems on certain benchmarks and delivered strong results in isiXhosa text generation despite its relatively small size.
The team stressed that MzansiLM is not designed as a chatbot or consumer-facing AI assistant, but rather as a foundation model that developers can adapt for specialised applications.
“In practice, that means developers could build tools for specific use cases; for example, summarising information or annotating raw data, in South African languages,” Meyer said.
“Adapting MzansiLM for a limited use case might be more effective and affordable than relying on proprietary large language models, if you want users to be able to interact with a system in their home language.”
The researchers said the work also highlights why even the world’s largest commercial AI systems still struggle to operate effectively in many African languages.
“Our findings show that the model can work well when fine-tuned for specific tasks but is not yet able to work well for general-purpose user interaction or instruction following, due to the limited training data,” Buys said.
“This helps to explain why even larger language models don’t yet work as well when used in languages other than English.”
The team said broader collaboration across Africa’s AI research community will be necessary to improve language coverage and expand AI accessibility for African language speakers.
“A lot of the progress we were able to make depends on earlier open research from the African Natural Language Processing research community, so continuing that openness is essential,” Lombard said.
“We still need better and broader data sources, stronger benchmarks, and the kind of shared datasets, models, code, and results that make it possible for others to reproduce and extend the work.”
Meyer said open collaboration remains critical to improving AI systems for African languages.
“The research community plays an important role here by working openly, sharing datasets, models, and findings so others can build on them. That kind of openness is often what leads to progress, especially compared to proprietary systems where much of the data and methodology isn’t accessible,” Meyer said.
The UCT research team has made both MzansiText and MzansiLM publicly available, with the paper titled MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages published on arXiv.








