Burmese · corpus in collection
မြန်မာဘာသာ
Burmese · 43M speakers · fine-tuning on a community-curated corpus.
The same thought, double the cost.
Same thought · 8 tokens in English, 17 in Burmese. Over 4× the cost to think in your own language.
A script that predates the printing press
Burmese has been written continuously for over a thousand years — palm-leaf manuscripts, parabaik folding books, monastic libraries whose catalogs pre-date the Domesday. The characters are round because a straight stroke would split a palm leaf; the whole visual grammar carries the memory of its substrate.
In modern LLM vocabularies, this thousand-year record is compressed into an average of 4.2 tokens per word. English gets 1.1. The gap is not linguistic — it is bookkeeping.
Zawgyi, and the cost of an encoding war
Until roughly 2019, most Burmese text on the internet was written in Zawgyi — a non-standard encoding that looked correct on screen but was gibberish to anything that spoke Unicode. A decade of blog posts, comment threads, and news archives sits in this format. To the tokenizer, they are two different languages.
Normalization is not a preprocessing step. It is a scribal act — an emendation, made once, then attested in the colophon.
Fine-tuning, not retraining.
We do not retrain from scratch. We do not need to. We start from an existing pretrained open model — Gemma-4, among the candidate bases under community review, its weights frozen — and add a small LoRA adapter shaped by community judgment: which texts matter, which spellings are contested, which voices the crawl missed.
The base model is given. What it's add is the language, in its own script, in the community's own hand.
We work with open source models. Gemma-4 (and SEA-LION )is being used for Burmese, Karen, and Jinghpaw as well as; Llama-3.1-8B may be used for Rohingya, for byte-level fallback where Hanifi is unrepresented.
Community participation and consent for curation and capture. Provenance recorded and fair compensation provided.
Converting legacy encodings and non-standard scripts into universal digital formats (e.g., Zawgyi to Unicode) is an act of cultural and historic preservation. Working alongside local linguists and community members, we systematically standardize text to be machine-readable while ensuring cultural context, orthographic nuances, and historical integrity remain intact.
Adapting foundational base models to local linguistic realities through targeted, data-efficient techniques calibrated to each language’s structure and context. Training logs and loss curves are published directly to the public archive with complete technical transparency.
Community-annotated benchmarks. Evals and benchmarks designed with local community and experts. Annotators acknowledgements and compensation. Inter-annotator agreement reported alongside every score.
- Q01Which pre-1990 orthographic conventions to preserve alongside the standard.
- Q02How to weight monastic Pali-inflected registers against contemporary speech.
The corpus belongs to the communities that made it. Consent is per item and revocable. Revocation is honored in the weights, not just in the archive.
Model weights reflect the terms and governance frameworks chosen with the community.