• Wed Jul 22 2026
Logo

Large Language Models for Small Language Communities



LONDON—Much has been said about AI’s security and military risks. But the cultural and economic threats posed by large language models (LLMs) trained on a limited number of languages and owned by a handful of multinational companies have been largely overlooked. The market dominance of these LLMs leaves developing countries with minority languages at a distinct disadvantage.

Ever since OpenAI released ChatGPT in November 2022, industry leaders and policymakers have touted AI’s potential to improve society by revolutionizing health care and education, as well as by driving productivity growth and income gains. But these benefits are not equally shared.

Many developing economies cannot access these AI systems, owing to their language biases. Most LLMs are trained on only 100 or so of the more than 7,000 languages spoken around the world, with English and the other Big Five “high-resource” languages—Chinese, Spanish, French, and German—accounting for nearly all of the data used. But only 400 million people are native English speakers, and just 25% of the global population speaks one of the Big Five as their first language.

This imbalance could lead to cultural and economic isolation in the short term and strangle growth and diminish linguistic diversity in the long term.

There is also the bitter irony that rich countries’ rapid adoption of AI, supported by a massive buildout of data centers equipped with ever more powerful compute infrastructure, has made it harder for users in developing countries to access the internet and benefit from digital tools. The AI-driven chip demand has raised the cost of smartphones, putting them out of reach for millions of people.

LLMs must better reflect local languages and cultures. Otherwise, communities risk losing their unique lexicons, long shaped by historical events, mass migrations, and political legacies. Moreover, the current lack of language diversity in AI systems will slow adoption. Digital financial services in Peru must communicate colloquially to be relevant to consumers. To deliver effective treatment, online health care in Vietnam requires the clarity that only a native language model can provide.

Mobile network operators (MNOs) are uniquely positioned to create “local” language models, owing to their armies of developers and data-processing capabilities.

They are also licensed and trusted government partners. This is important because a major barrier to creating LLMs for minority languages is the lack of training data. The solution lies in the huge amounts of data held by governments, which tend to be sensitive and confidential.

For example, Ukraine’s largest MNO, Kyivstar, created a partnership with the Ministry of Digital Transformation to develop the country’s first national LLM, which was tailored to Ukrainian as well as other languages spoken in the country, including Russian, Bulgarian, and Crimean Tatar. Complicating the task was the fact that much of the literature in Ukraine was written under Russian rule and thus considered culturally inappropriate. As a result, four advisory committees with binding authority over the model’s technical, legal, cultural, historical, and linguistic aspects guided the curation of training texts.

The Global System for Mobile Communications Association (GSMA), of which I am director general, is pursuing a similar strategy in much of Sub-Saharan Africa, including Nigeria, Madagascar, Togo, and the Democratic Republic of the Congo, through the African AI Language Models project. The initiative is bringing together the continent’s Group of Six MNOs—Airtel, Axian Telecom, Ethio Telecom, MTN, Orange, and Vodacom—and other telecommunications companies, public and private actors, academics, and civil-society groups to build scalable and inclusive LLMs for African languages.

To scale AI startups that generate positive socioeconomic outcomes for the underserved, the GSMA has launched the EmergingTech program, funded by the United Kingdom’s Foreign, Commonwealth, and Development Office. Its first cohort includes ToumAI, which creates voice interfaces in Morocco’s local dialects to help low-literate, rural users access digital and financial services through simple spoken commands, and Wiseyak, which is developing a multilingual mix-code conversational AI solution to help reduce Nepal’s digital divide.

These initiatives are extremely important as the world moves from prompt-driven to agentic AI. Digital services such as banking, health care, and education are increasingly embedding AI to assist users and will become crucial to supporting lives and livelihoods. Firms, too, are benefiting from open-source models.

Investment in local AI creates a virtuous circle for developing economies, generating a pool of tech talent that is critical to digital independence. As concerns over data and data-processing sovereignty grow, these economies must improve operational control and governance over the deployment of agentic AI. That means collaborating with developers and MNOs to build minority language models that meet their society’s needs.

Vivek Badrinath is Director General of the Global System for Mobile Communications Association.

Copyright: Project Syndicate, 2026.
www.project-syndicate.org