A community language-data infrastructure
A digital future for underrepresented languages.
DR Congo Starting with Lingala and the Congolese context.
Lokota helps Congolese communities document, validate, structure and govern their language resources — text, translations, varieties, voice — with an approach designed to be adapted to other underrepresented languages.
Your language isn’t on Lokota yet? Propose a language

Our languages, our voice, our digital future.
Why Lokota
Languages spoken by millions of people stay almost invisible in the digital world.
Without reliable data, a language is missing from search, translation tools, speech recognition and AI systems. That limits access to information, education and services — and a language absent from today’s data risks being absent from tomorrow’s technology.
What guides Lokota
Community first
Speakers contribute to, validate and govern their own resources.
Consent and ethics
Clear data use, attribution, transparent reuse conditions.
Quality and diversity
Not just more data: reliable resources that represent real-world usage.
Responsible reuse
Governed dataset access for research, education and more inclusive AI.
Contribute
Three ways to grow a language
Anyone can take part, whatever their expertise. The first contribution stays simple; advanced metadata comes later.
Languages
- Words and expressions
- Translations
- Regional and urban varieties
- Contextual usage examples
Voice
- Pronunciations
- Audio recordings
- Authentic speaker audio
- Oral knowledge
Data
- Structured resources and metadata
- Community validation
- Open datasets
- Reuse for research and AI
How Lokota works
From contribution to the tools of tomorrow
Lokota is not just a dictionary: it is a pipeline that turns contributions into reusable data infrastructure.
Collect
Communities submit words, phrases, translations, varieties and voice.
Verify
Contributions are independently reviewed and confirmed.
Structure
Information becomes documented, standardised data.
Share
Datasets and resources made available for research and education.
Build
Translation, voice, learning tools, multilingual assistants, research.
From language to AI
One word, documented properly, opens many possibilities
The application layer comes after reliable data is created — never before.
Our first implementation
Starting with the languages of the Democratic Republic of the Congo
Lingala is the pilot language. Kikongo, Tshiluba and Kiswahili follow, then priority regional languages. The method and tooling are built to be adapted, later, to other underrepresented languages.

Languages on Lokota
The Lokota method can later be offered to other underrepresented languages beyond the DRC.
A community, not a database
Every voice counts, every skill counts

Speakers
Share their language, their words and their knowledge.

Checkers
Independently review and validate contributions.

Community leaders
Guide priorities and governance for their language.

Linguists
Bring expertise on varieties and usage.

Researchers
Reuse the data for reproducible research.

Partners
Institutions, NGOs and universities that support and collaborate.
Governance and trust
Governance is a feature, not a legal page.
Every contribution keeps a record, a consent basis and clear reuse conditions.
- Community validation
- Informed consent
- Attribution
- Transparent contribution history
- Recognition of language varieties
- Ability to contest a classification
- Independent checking
- Responsible dataset access
Research and institutions
An open language-data infrastructure
Lokota is not only a collection platform: it is a community infrastructure that can support reproducible research and responsible technology development.
Resources
Go straight to Lokota’s content
Datasets
Structured, open resources with reuse conditions.
Publications
Research work and methodological notes.
Projects
Collection campaigns and ongoing collaborations.
Events
Meetups, workshops and talks.
Documentation
Contribution and validation guides.
Governance
Policies, consent and access conditions.
Join the movement
Contribute, explore, collaborate.
A community-governed infrastructure so underrepresented languages have the reliable data they need — beginning with Lingala.
