Rohingya · corpus in collection
𐴌𐴟𐴇𐴥𐴝𐴚𐴒𐴠𐴝
Rohingya · ≈2.8M speakers · fine-tuning on a community-curated corpus.
The same thought, double the cost.
Hanifi shatters into raw bytes. The task is prior to fluency: teach the machine these bytes are letters.
A script the vocabulary has not met
Hanifi Rohingya entered Unicode in 2018. Almost every large model's tokenizer was trained before then. To those models, a Rohingya word is not rare — it is invisible: each letter shatters into three or four raw bytes.
Our work involves curating a clean Hanifi corpus (eg, Hanifi) and teaching the machine that these bytes are letters at all.
Three scripts, one language
Rohingya has been written in Hanifi, in Arabic script, and in Latin transliteration — each carrying its own politics, its own community. We work in Hanifi here because the community we work with elected it. The others are named in the scholia; they are not averaged away.
This is not neutrality. It is a choice, credited to the people who made it.
Faces are not the point
Speakers appear as voice-lines, not portraits. This is not a stylistic preference. It is a deficiency. Some contexts require strict anonymity for the safety of the contributors.
We work with open source models. Gemma-4 (and SEA-LION )is being used for Burmese, Karen, and Jinghpaw as well as; Llama-3.1-8B may be used for Rohingya, for byte-level fallback where Hanifi is unrepresented.
Community participation and consent for curation and capture. Provenance recorded and fair compensation provided.
Converting legacy encodings and non-standard scripts into universal digital formats (e.g., Zawgyi to Unicode) is an act of cultural and historic preservation. Working alongside local linguists and community members, we systematically standardize text to be machine-readable while ensuring cultural context, orthographic nuances, and historical integrity remain intact.
Adapting foundational base models to local linguistic realities through targeted, data-efficient techniques calibrated to each language’s structure and context. Training logs and loss curves are published directly to the public archive with complete technical transparency.
Community-annotated benchmarks. Evals and benchmarks designed with local community and experts. Annotators acknowledgements and compensation. Inter-annotator agreement reported alongside every score.
- Q01Which of Hanifi, Arabic, and Latin scripts to serve first — and how the others follow.
- Q02Whether model weights should be publicly released, or held under community license.
The corpus belongs to the communities that made it. Consent is per item and revocable. Revocation is honored in the weights, not just in the archive.
Model weights reflect the terms and governance frameworks chosen with the community.