Minority language speakers have good reason to be wary of artificial intelligence. The chatbots that now draft emails and translate menus work best in the handful of languages that dominate the internet, and they stumble in languages spoken by a few thousand people. Yet a new account from two University of Melbourne researchers argues that the same technology is quietly handing those communities something they rarely had before: the ability to build their own language tools, on their own terms and at almost no cost.

In an article published by The Conversation on 29 September 2026 and republished by Tech Xplore, computational linguists Ekaterina Vylomova and Raphael Merx describe three projects. One documents Hula, a language of about 10,000 people in Papua New Guinea. One helps health educators translate training material into Tetun in Timor-Leste. The third is a Dinka–English dictionary app built by a single developer from South Sudan.

This article explains why AI leaves minority language speakers behind, how each of the three tools works, what they have produced and what they cost, where the optimism needs caution, and what developers, funders and businesses can take from them. For wider context on AI translation, see our coverage of Meta’s SeamlessM4T translation model and of Prime Video’s AI lip-sync for dubbed audio.

Why AI Leaves Minority Language Speakers Behind

minority language ai communities lifeline b open book on stand

Every generative AI chatbot runs on a large language model, and that model learns from text. The more text a language has online, the better a model tends to handle it. Vylomova and Merx describe “a great divide between which languages get included, and which do not”. Majority languages dominate the training mix, while a minority language, one with very little representation online, may barely register at all.

Data decides who gets served

Ethnologue, the reference catalogue of the world’s languages, counts 7,170 languages in use today. Only a small fraction of them have the volume of digitised text, parallel translations and speech recordings that modern AI systems learn from. Researchers in natural language processing have long described most of the rest as low-resource languages. For them, the problem is circular. Without data there is no good tool, and without a good tool there is little reason for speakers to create data in their own minority language online.

Two languages at the edge

The researchers use two examples to show the scale of the gap. Hula, also known as Vula’a, is spoken in Papua New Guinea’s Central Province and has around 10,000 speakers. Tetun, the lingua franca of Timor-Leste, has just over 1 million. Both, they write, “represent a tiny fraction of AI model training data, and as such are very likely to be misrepresented in the output, if they show up at all”.

Speakers of the three languages in the article, as a share of Dinka’s roughly 5 million
Dinka, South Sudan, about 5,000,000 100%
Tetun, Timor-Leste, just over 1,000,000 20%
Hula, Papua New Guinea, about 10,000 0.2%

Why a bad translation matters

A clumsy translation of a shopping query is an annoyance. A clumsy translation of a health leaflet, a disaster warning or an official form is a risk, because the reader has no way to know which parts are wrong. That is the core worry for any minority language community: the tools that claim to support their language can do quiet damage when they guess.

There is an economic side too. The authors point to wider warnings that AI will reward capital owners at the expense of workers, and that technologically advanced countries will pull ahead of low-income ones, a pattern the UN Development Programme has called the “next great divergence”. Linguistic inequality, they suggest, could be one more layer of that divide, with every minority language on the wrong side of it.

Vavanagi: A Minority Language Platform the Hula Community Built Itself

minority language ai communities lifeline c studio microphone on stand

The most striking of the three minority language projects is Vavanagi, a language documentation platform for Hula. It was “built entirely by and for the Hula community, with coding help from AI”, according to The Conversation. The work was led by Bri Olewale, a community member who contributed to the article and is the first author of a paper describing the platform, posted to arXiv in March 2026 and revised in June.

How it started

The Hula community is organised online around a WhatsApp group with interest channels for canoe racing, church, history and language. In the language channel, members were discussing long-running Bible translation efforts and whether AI could speed them up. That raised an obvious need: a corpus of Hula text to work from. Olewale, who has a technical background but is not a web developer, set out to crowdsource one.

How the platform works

Vavanagi runs a four-stage pipeline. An administrator imports English source sentences. Community translators submit Hula translations as text and, where they can, as voice recordings. A small team of reviewers checks each entry for clarity, cultural appropriateness and consistency, and leaves comments rather than simply rejecting entries. Finally, the administrator exports approved records for archiving and research.

The roles map onto the community’s existing structure. The reviewers are the administrators of the WhatsApp language group, and the translators are its members. Voice capture was added so that elders who are less comfortable writing Hula could take part, and a leaderboard adds some friendly competition to this minority language effort.

Built with AI coding help for under $20

The technical side is deliberately modest. The paper says early development started in Replit, a browser-based coding tool, before moving to a code editor as the platform matured. Records live in Firebase Firestore, a cloud document database that also handles sign-in and role-based access. Total project technology costs, including the domain name, have been “less than $20 USD”.

That figure is the heart of the argument for minority language communities. Normally, building a documentation platform for a minority language needs linguists and software engineers the community does not have. Here, AI coding assistants filled the engineering gap, while the language expertise stayed with the speakers.

Paying translators in kina

The community also designed its own way to fund the work. Urban Hula speakers, who may use the language less day to day but have more disposable income, pay into a shared prize pool that rewards village-based translators. The paper puts the current incentive at PGK0.10 per sentence, funded by voluntary contributions of PGK10 to PGK100 from members who do not translate.

What it has produced

Uptake was fast. The first batch of 2,000 sentences was finished in two weeks and a second batch of 1,500 in three days. Sentences are short, with a median of eight words, and 91% of submissions were approved by reviewers on the first pass.

Vavanagi measureFigureSource
Hula speakersAbout 10,000Paper, citing Ethnologue
Translators and reviewers77 translators, 4 reviewers, 1 administratorPaper
Parallel English–Hula sentences12,124, covering about 9,000 unique Hula wordsPaper
First-pass approval rate91%Paper
Usability score (System Usability Scale)73.4, from 8 translatorsPaper
Technology costs to dateUnder US$20, including the domainPaper
Vavanagi translation pace, sentences per day (batch size divided by days taken)
Batch 1: 2,000 sentences in 14 days about 143 a day
Batch 2: 1,500 sentences in 3 days 500 a day

Reviewers also noticed something the corpus captures by accident: the language changing in real time. Translations include loanwords from Tok Pisin, Papua New Guinea’s main lingua franca, such as “ketolo” for kettle and “lamepa” for lamp, along with frequent code-switching. Adolescents in the village often speak Tok Pisin to each other while still understanding their parents’ Hula, a pattern that puts any minority language under pressure.

What comes next

The goal, according to The Conversation, is to collect enough data to build a Hula translator app. The paper sets out the order: a machine translation model first, then automatic speech recognition, leading to a voice-enabled Hula–English app for everyday market trade and media. The team also plans a voice-only mode in which elders speak and younger members transcribe, splitting the work between fluency and digital literacy.

Tulun: Minority Language Translation for Health Workers in Timor-Leste

minority language ai communities lifeline d lifebuoy ring

The second tool, Tulun, tackles a different problem. Timor-Leste’s health educators need training material in Tetun, and general-purpose machine translation handles specialised medical vocabulary badly. Tulun was built in partnership with Maluk Timor, a local non-governmental organisation, and lets its staff manage their own list of approved terms and phrases.

Merx, one of the article’s two authors, developed Tulun as part of his research, a point The Conversation discloses. The platform was presented at ACL 2025, one of the main computational linguistics conferences, in a system-demonstration paper by Merx, Hanna Suominen, Lois Yinghui Hong, Nick Thieberger, Trevor Cohn and Vylomova.

Glossaries instead of retraining

Most ways of adapting machine translation to a specialist field involve fine-tuning a model, which is impractical for small organisations with no machine learning staff. Tulun takes a different route. It combines a neural translation model with post-editing by a large language model, guided by glossaries and translation memories the users maintain themselves. Health experts decide how a term should be rendered, and the AI applies their decisions.

That turns translation into what the researchers call a collaborative exercise between AI and Tetun health experts. For a minority language with little specialist text online, it is a practical way to get accuracy without waiting for a better model.

What the evaluations show

The ACL paper reports sizeable gains. On medical and disaster-relief translation tasks for Tetun and Bislama, a language of Vanuatu, Tulun improved on baseline machine translation systems by 16.90 to 22.41 ChrF++ points. ChrF++ is a standard translation quality score that compares character and word sequences with a reference translation, so a gain of that size is large.

Across six low-resource languages on FLORES, a widely used benchmark, Tulun beat both standalone machine translation and plain large language model approaches, with an average improvement of 2.8 ChrF++ points over NLLB-54B, Meta’s large multilingual translation model. The platform is open source, which matters for any minority language group that wants to run its own copy.

The Dinka–English Dictionary: One Developer, Five Million Speakers

minority language ai communities lifeline e globe on stand

The third minority language example is the smallest in team size and the largest in potential audience. Dinka has about 5 million speakers across South Sudan, yet it has very limited digital representation. Alier Makoi Achuoth, a developer from South Sudan, built a Dinka–English dictionary app that lets Dinka speakers translate words and verify their meanings, so the vocabulary grows through community checking.

AI models helped Achuoth design and build the app, and collect and organise its data. He told the researchers he wanted to “develop a practical solution instead of waiting for another person or organisation to address the problem”. The app’s Google Play listing describes an offline dictionary with search, bookmarks and word definitions, showed more than 1,000 downloads and was last updated on 8 August 2026 when we checked.

That download count is modest next to 5 million speakers, and it would be wrong to overstate the app’s reach. But its existence makes the article’s point: a single person with AI assistance produced a working minority language resource that no company or agency had delivered.

What the Three Minority Language Tools Have in Common

minority language ai communities lifeline f portable radio

The three projects serve different purposes, but Vylomova and Merx argue they share one pattern: each uses AI “to better channel local knowledge back within their language community”. The table below sets them side by side.

ToolLanguage and speakersWho built itWhat AI didPurpose
VavanagiHula, Papua New Guinea, about 10,000Hula community, led by Bri OlewaleCoding help to build the platformCrowdsourced text and voice corpus, elder-led review
TulunTetun, Timor-Leste, just over 1 millionRaphael Merx with Maluk TimorTranslation plus model post-editing guided by glossariesHealth education material for health workers
Dinka–English dictionaryDinka, South Sudan, about 5 millionAlier Makoi AchuothDesign, build, data collection and organisationVocabulary, translation and word verification

AI as a builder’s tool, not the product

In none of the three cases is a chatbot the thing people use. AI sits one step back, writing code, organising data or post-editing a draft, while the finished tool is shaped by speakers. That distinction matters for any minority language, because it keeps the judgement about what is correct with the people who actually speak it.

Local knowledge stays in charge

For Vavanagi, that means elders reviewing a language that is losing ground to Tok Pisin among Hula youth. For Tulun, it means Timorese health educators adapting material for Timorese health workers. For the Dinka dictionary, it means Dinka speakers verifying the meanings of their own words.

Costs low enough to skip the grant

The researchers argue that AI is reducing the time and money needed to build these tools “to the point where communities can own and build them without any external funding”. They link this to a broader shift. AI adoption in Kenya and Nigeria, they note, is as high as in the United States according to the World Bank, which undercuts the idea of low-income countries as AI laggards. They add that small firms in the Global South can now afford machine translation of marketing content they could never have paid a professional to translate.

Five Levels of Community Control

The Vavanagi paper adds a useful piece of vocabulary: a five-level scale for how much control a community really has over a language technology project. Many projects describe themselves as “community-based”, the authors note, which hides whether the community was consulted once or runs the whole thing.

LevelNameWho holds control
1Community consultedMembers are consulted, but day-to-day activity stays external
2Community engaged, externally ledMembers contribute data or feedback; planning and operations stay external
3Community-led operations, externally designedMembers run daily work; core design comes from outside
4Community-led with co-shaped designStarted externally, but community priorities reshape goals and methods
5Fully community-initiated and governedInitiative, decisions, implementation and data governance all sit within the community

The authors place Vavanagi at Level 5 and call it, to their knowledge, the first community-led language technology initiative for a language of its size. Other community-led efforts, such as the Masakhane network for African languages, work on languages with millions of speakers, at least two orders of magnitude more than Hula. The approach also lines up with the CARE principles for Indigenous data governance: collective benefit, authority to control, responsibility and ethics.

Where the Minority Language Optimism Needs Caution

The article is a hopeful one, and deliberately so. It is still worth separating what the evidence shows from what it suggests.

The big models still fall short

Nothing in these projects fixes the underlying gap. The authors say so themselves: AI models “might continue to poorly incorporate minority languages”. Their bet is that community tools will “become the best representation of minority languages online”, which is a different outcome from mainstream chatbots becoming good at Hula or Dinka.

Small numbers and involved authors

The evidence base is early. The Vavanagi usability score comes from eight translators, and its corpus figures are self-reported by a team that includes the project’s own leader. Merx developed Tulun, and the Tulun evaluation is his group’s own. None of that makes the results wrong, and the disclosures are clear, but independent evaluation would strengthen the case for every minority language tool described here.

Who controls the data in practice

Level 5 governance describes who decides, not where data physically lives. Vavanagi runs on Firebase, a commercial cloud service. A corpus of 12,124 parallel sentences in a rare language could become valuable to AI developers, so the community’s rules on licensing and reuse will matter as much as its platform design. The paper frames data sovereignty as central, and the next test is how it handles requests to use the data.

Keeping tools alive

Low build costs do not remove maintenance costs. Apps need updates, domains need renewing and volunteers move on. Projects that depend on one person, like the Dinka dictionary, are especially exposed, which is worth remembering before treating near-zero development costs as a complete answer.

What Developers, Funders and Businesses Can Learn

The lessons reach beyond language preservation. They say something about how AI is changing who can build software, and how organisations should work with communities rather than simply collecting their data.

For AI developers

The Tulun result is a reminder that a glossary-guided pipeline can beat a far larger general model on specialist text. For teams serving any minority language market, investing in terminology control and human review may deliver more than waiting for the next model release. Meta’s NLLB and SeamlessM4T projects made broad coverage possible; tools like Tulun show how to make it accurate in a narrow domain.

For researchers and funders

Vylomova and Merx argue for moving away from top-down research towards community-led projects, especially with Indigenous communities in Australia. They quote Cat Kutay, a computer scientist of Aboriginal descent at Charles Darwin University, who says several First Nations are building their own language technology: “Like we took up the motorcar when we were shown that, as our life evolves around travel, so we are taking up AI as our life evolves around translations and storytelling.”

For businesses that work across languages

Businesses face a smaller version of the same problem whenever they translate product names, safety instructions or contracts. The practical lessons carry over: keep an approved glossary, let people who know the subject review output, and treat AI translation as a draft. Our AI models, tools and releases hub tracks the models that sit behind these workflows, and our look at how VidMage sells in 25 languages shows what goes wrong when localisation stops at the interface.

When Vavanagi was featured on Papua New Guinea’s NBC Radio, Hula elder Alu Rigo Ravu Siro described it the way her people describe a voyage. Hula canoes are double-hulled and need a crew: some sail, some bail, some cook, some mind the children. “Just come, join in,” she said. It is as good a summary as any of how a minority language community can make AI work for it.

Minority Language AI FAQ

Why does AI perform badly in minority languages?

AI language systems learn from text, and a minority language usually has very little of it online. With so little training data, models misrepresent these languages in their output, or fail to handle them at all.

What is Vavanagi?

Vavanagi is a community-run platform for documenting Hula, a language of about 10,000 people in Papua New Guinea. Translators submit English–Hula translations in text and voice, and elders review them. It has produced 12,124 sentence pairs for under US$20 in technology costs.

What is Tulun?

Tulun is an open-source translation platform that combines machine translation with large language model post-editing guided by user-managed glossaries. It is used by health educators at Maluk Timor to translate material into Tetun.

Can small communities build language apps without funding?

The three projects suggest they increasingly can, because AI coding assistants reduce the engineering work. Maintenance, hosting and data governance still need planning, so near-zero build costs are a starting point rather than the whole budget.

Will large AI models eventually fix minority language support on their own?

The researchers are sceptical. They expect big models may continue to handle these languages poorly, and argue that tools built by the communities themselves are likely to become the best representation of each minority language online.

References