Rohingya · corpus in collection

𐴌𐴟𐴇𐴥𐴝𐴚𐴒𐴠𐴝

Rohingya · ≈2.8M speakers · fine-tuning on a community-curated corpus.

I COULD BE BOUNDED IN A NUTSHELL AND COUNT MYSELF A KING OF INFINITE SPACE. [ 2.2.254-255 ]𐴌𐴟𐴇𐴥𐴝𐴚𐴒𐴠𐴝 · Rohingya in Bangladesh, Myanmar, and diaspora

The same thought, double the cost.

ENGLISH · EN-US
The river remembers what the map forgets.
words
8
tokens
8
tok/word
1.0
ROHINGYA · ROHINGYA
<0xF0><0x90><0xB4><0x93><0xF0><0x90><0xB4><0xA0><0xF0><0x90><0xB4><0x9F><0xF0><0x90><0xB4><0x9A><0xF0><0x90><0xB4><0x8C>
words
3
tokens
20
tok/word
6.7
1.0 tok/wordvs6.7 tok/word· 2.5× the cost to think in your own language

Hanifi shatters into raw bytes. The task is prior to fluency: teach the machine these bytes are letters.

A script the vocabulary has not met

Hanifi Rohingya entered Unicode in 2018. Almost every large model's tokenizer was trained before then. To those models, a Rohingya word is not rare it is invisible: each letter shatters into three or four raw bytes.

Our work involves curating a clean Hanifi corpus (eg, Hanifi) and teaching the machine that these bytes are letters at all.

Three scripts, one language

Rohingya has been written in Hanifi, in Arabic script, and in Latin transliteration each carrying its own politics, its own community. We work in Hanifi here because the community we work with elected it. The others are named in the scholia; they are not averaged away.

This is not neutrality. It is a choice, credited to the people who made it.

Faces are not the point

Speakers appear as voice-lines, not portraits. This is not a stylistic preference. It is a deficiency. Some contexts require strict anonymity for the safety of the contributors.

section numbers · toggle blueprint for configs
[00] · plate

We work with open source models. Gemma-4 (and SEA-LION )is being used for Burmese, Karen, and Jinghpaw as well as; Llama-3.1-8B may be used for Rohingya, for byte-level fallback where Hanifi is unrepresented.

Curation
[01] · plate

Community participation and consent for curation and capture. Provenance recorded and fair compensation provided.

Normalization
[02] · plate

Converting legacy encodings and non-standard scripts into universal digital formats (e.g., Zawgyi to Unicode) is an act of cultural and historic preservation. Working alongside local linguists and community members, we systematically standardize text to be machine-readable while ensuring cultural context, orthographic nuances, and historical integrity remain intact.

[03] · plate

Adapting foundational base models to local linguistic realities through targeted, data-efficient techniques calibrated to each language’s structure and context. Training logs and loss curves are published directly to the public archive with complete technical transparency.

[04] · plate

Community-annotated benchmarks. Evals and benchmarks designed with local community and experts. Annotators acknowledgements and compensation. Inter-annotator agreement reported alongside every score.

THE CONTESTED QUESTIONS · KEPT VISIBLE, NOT RESOLVED FOR THE COMMUNITY
  • Q01Which of Hanifi, Arabic, and Latin scripts to serve first — and how the others follow.
  • Q02Whether model weights should be publicly released, or held under community license.
[ 06 ] THE VOICES
named with consent · faces omitted by choice
anonymous
anonymous
teacher, Cox's Bazar
duration · 1:12
anonymous
anonymous
poet
duration · 2:04
Mohammed A.
named
diaspora journalist
duration · 1:48
[ 06 ] SOVEREIGNTY

The corpus belongs to the communities that made it. Consent is per item and revocable. Revocation is honored in the weights, not just in the archive.

Model weights reflect the terms and governance frameworks chosen with the community.

[ 08 ] COLOPHON

Completed at location withheld, in 2026, on the Rohingya project, rank TBD. Released with community consent; evaluation held.

place
location withheld
date
2026 — provisional
hardware
TBD
license
under community deliberation
elsinore data co · undertext · v0.5 · 2026.07
next projectVanuatu cluster
contribute