Infinite Space Language Project
Culture is grounded
in the language it speaks.
Language has always encoded culture; now, with AI, it encodes how machines see it too. We fine-tune open LLMs on corpora curated by the communities themselves — so no language is left out of the intelligence revolution to come.
43M speakers.
≈2M speakers.
≈1.5M speakers.
≈2.8M speakers.
≈340k across 138 languages speakers.
If your language, community, or archive belongs on this list, reach out.
We work with open source models. Gemma-4 and SEA-Lion are being used for Burmese, Karen, and Jinghpaw; Llama-3.1-8B may be used for Rohingya, for byte-level fallback where Hanifi is unrepresented.
Community participation and consent for curation and capture. Provenance recorded and fair compensation provided.
Converting legacy encodings and non-standard scripts into universal digital formats (e.g., Zawgyi to Unicode) is an act of cultural and historical preservation. Working alongside local linguists and community members, we systematically standardize text to be machine-readable while ensuring cultural context, orthographic nuances, and historical integrity remain intact.
Adapting foundational base models to local linguistic realities through targeted, data-efficient techniques calibrated to each language’s structure and context. Training logs and loss curves are published directly to the public archive with complete technical transparency.
Community-annotated benchmarks. Evals and benchmarks designed with local community and experts. Annotators acknowledgements and compensation. Inter-annotator agreement reported alongside every score.
The same thought, double the cost.
Same thought · 8 tokens in English, 17 in Burmese. Over 4× the cost to think in your own language.
Adopter-agnostic, representation-dependent.
Large language models are adopter-agnostic and representation-dependent: anyone can use them, but they only serve what they have seen — and what they have seen more so follows money, not meaning. For a low-resource language, that produces two compounding harms.
The first is a tax. The same sentence costs far more tokens in an under-represented script than in English. More tokens means more corpus, more compute, worse latency, shorter effective context — a structural penalty for thinking in your own language.
The second is a deficiency. Some scripts — Hanifi Rohingya, added to Unicode in 2018 — are effectively absent from model vocabularies. Each letter shatters into raw bytes. Here the task is prior to fluency: teach the machine that those bytes are letters at all.
Fine-tuning inverts the economics.
Fine-tuning doesn't need pre-training-scale resources. It needs the one thing the big corpus never had: community judgment about which texts matter, which spellings are contested, which voices are missing. That judgment is the scarce input, and it's a resource these communities have in abundance.
The corpus belongs to the communities that made it. Consent is per item and revocable. Revocation is honored in the weights, not just in the archive.
Model weights reflect the terms and governance frameworks chosen with the community.
Some of this corpus, however, can only exist if no one can prove who made it.
In conflict, high-risk, or humanitarian contexts, strict operational security and total contributor anonymity are the necessary prerequisite for the data existing at all. Operating in these regions requires trusted local partnerships, deep field experience, and physical presence where others cannot go.
Privacy-enhancing technologies, developed with experts, so a contribution can be learned from without the raw data or the person ever being exposed.
Industry-standard protection for data at rest and in transit, and — because of where we work — operational security for people living under authoritarian or occupying regimes.
Were a phone searched at a checkpoint, nothing on it would expose their contributions here.