Vanuatu local languages · corpus in collection

Vanuatu local languages · ≈340k across 138 languages speakers · fine-tuning on a community-curated corpus.

I COULD BE BOUNDED IN A NUTSHELL AND COUNT MYSELF A KING OF INFINITE SPACE. [ 2.2.254-255 ] · under co-design with local partners

The same thought, double the cost.

ENGLISH · EN-US
The river remembers what the map forgets.
words
8
tokens
8
tok/word
1.0
VANUATU LOCAL LANGUAGES · VANUATU
Wota i tingting long samting we map i fogetem.
words
8
tokens
14
tok/word
1.8
1.0 tok/wordvs1.8 tok/word· 1.8× the cost to think in your own language

Bislama (creole) shown here as a sample; individual Vanuatu languages will each have their own tokenizer twin.

The world's densest language geography

Vanuatu has more languages per capita than any country on Earth 138 languages for roughly 340,000 speakers. Some are documented; many are not. Some communities want their languages modeled; others do not.

This page is a cluster hub. It says so plainly. The base model itself is dashed on this page: under evaluation, pre-selection.

Pre-selection is not indecision

Choosing which languages to work on is the first act of the project, and it belongs to local partners not to us.

Provisional honesty is the design.

section numbers · toggle blueprint for configs
[00] · plate

We work with open source models. Gemma-4 (and SEA-LION )is being used for Burmese, Karen, and Jinghpaw as well as; Llama-3.1-8B may be used for Rohingya, for byte-level fallback where Hanifi is unrepresented.

Curation
[01] · plate

Community participation and consent for curation and capture. Provenance recorded and fair compensation provided.

Normalization
[02] · plate

Converting legacy encodings and non-standard scripts into universal digital formats (e.g., Zawgyi to Unicode) is an act of cultural and historic preservation. Working alongside local linguists and community members, we systematically standardize text to be machine-readable while ensuring cultural context, orthographic nuances, and historical integrity remain intact.

[03] · plate

Adapting foundational base models to local linguistic realities through targeted, data-efficient techniques calibrated to each language’s structure and context. Training logs and loss curves are published directly to the public archive with complete technical transparency.

[04] · plate

Community-annotated benchmarks. Evals and benchmarks designed with local community and experts. Annotators acknowledgements and compensation. Inter-annotator agreement reported alongside every score.

THE CONTESTED QUESTIONS · KEPT VISIBLE, NOT RESOLVED FOR THE COMMUNITY
  • Q01Which languages within the cluster to serve first, and by whose invitation.
  • Q02How to model without extracting — the whole question.
[ 06 ] THE VOICES
named with consent · faces omitted by choice
not yet recorded
anonymous
duration ·
[ 06 ] SOVEREIGNTY

The corpus belongs to the communities that made it. Consent is per item and revocable. Revocation is honored in the weights, not just in the archive.

Model weights reflect the terms and governance frameworks chosen with the community.

[ 08 ] COLOPHON

Completed at TBD, in 2026, on the Vanuatu local languages project, rank TBD. Released with community consent; evaluation held.

place
TBD
date
2026 — provisional
hardware
TBD
license
TBD
elsinore data co · undertext · v0.5 · 2026.07
next projectBurmese
contribute