Every language deserves working technology.
We build and maintain free, open language technology for the thousands of languages that commercial systems ignore — and we publish an honest, inspectable record of what does and doesn't exist for each of them.
Five pillars
Identify
Know which language and script a text is in — the first step for any corpus.
Collect
Find and clean web text for languages that big crawls miss.
Model
Train and release multilingual models that include the long tail.
Evaluate
Benchmark systems per language and per script so gaps are visible.
Document
Keep a public directory of what technology exists for each language.
Over 7,000 languages are spoken today, yet the web corpora, language models and OCR engines that power modern NLP cover only a small fraction of them. When a language can't be identified, it can't be collected; when it can't be collected, no model learns it; when no model learns it, its speakers are left out.
CommonGlot breaks that chain at each link. GlotLID and GlotScript identify languages and writing systems. GlotCC, GlotWeb and our smaller corpora collect text. Glot500 turns it into models. GlotOCR Bench measures how well today's systems read the world's scripts. Everything is released under open licences.
Alongside the tools, we maintain a directory that is part catalogue and part audit: one booklet per language and per script, built from public sources, showing exactly which technologies support it.
Four kinds of open deliverables
Language & script identification
Models and libraries that tell you what language and writing system a piece of text is in — for 2,000+ language labels and every Unicode script.
- fastText LID model (v1–v3)
- Python script detector on PyPI
- Language→script resource for 8,000+ languages
Web-scale and curated corpora
Clean, documented text for minority languages: a CommonCrawl corpus for 1,000+ languages, a web index for 400+, and smaller curated collections.
- GlotCC-V1 on Hugging Face
- 169,155+ verified web links
- Storybooks in 180 languages
Multilingual language models
Pretrained encoders that extend coverage from about a hundred languages to more than five hundred.
- Glot500-m (XLM-R-base extended)
- Glot500-c training corpus
- Head vs. tail evaluation suite
Benchmarks and leaderboards
Evaluation that is broken down by script and language, so you can see who is served and who is not.
- OCR benchmark across 150+ scripts
- Results for 14 OCR models
- Public leaderboard
The language & script directory
A booklet per language and per script, combining Glottolog metadata with GlotSuite coverage. Browse it on this site or download it as JSON.
- Language booklets
- Script booklets
- Downloadable JSON
Principles
Open by default
Code, models and data packaging are released under permissive licences (Apache-2.0, MIT, CC0) wherever the source allows.
Every claim has a source
Directory entries are generated from Glottolog, GlotScript, LinguaMeta and our own releases, and link back to them.
Coverage over headlines
We report per-language and per-script results, not just averages that hide the long tail.
Communities first
Speakers know their languages best. Corrections and contributions from communities take priority.