Language tech for the long tail.
Most language technology serves a few dozen languages. CommonGlot builds and documents the open tools, corpora and benchmarks that reach the rest — and keeps a public record of what exists for every language and every script.
The directory in numbers
Every language deserves working technology.
We build and maintain free, open language technology for the thousands of languages that commercial systems ignore — and we publish an honest, inspectable record of what does and doesn't exist for each of them.
Read the mission →Identify
Know which language and script a text is in — the first step for any corpus.
Collect
Find and clean web text for languages that big crawls miss.
Model
Train and release multilingual models that include the long tail.
Evaluate
Benchmark systems per language and per script so gaps are visible.
Four kinds of open deliverables
Language & script identification
Models and libraries that tell you what language and writing system a piece of text is in — for 2,000+ language labels and every Unicode script.
- fastText LID model (v1–v3)
- Python script detector on PyPI
- Language→script resource for 8,000+ languages
Web-scale and curated corpora
Clean, documented text for minority languages: a CommonCrawl corpus for 1,000+ languages, a web index for 400+, and smaller curated collections.
- GlotCC-V1 on Hugging Face
- 169,155+ verified web links
- Storybooks in 180 languages
Multilingual language models
Pretrained encoders that extend coverage from about a hundred languages to more than five hundred.
- Glot500-m (XLM-R-base extended)
- Glot500-c training corpus
- Head vs. tail evaluation suite
Benchmarks and leaderboards
Evaluation that is broken down by script and language, so you can see who is served and who is not.
- OCR benchmark across 150+ scripts
- Results for 14 OCR models
- Public leaderboard
The language & script directory
A booklet per language and per script, combining Glottolog metadata with GlotSuite coverage. Browse it on this site or download it as JSON.
- Language booklets
- Script booklets
- Downloadable JSON
A booklet for every language and script
Like Glottolog or Ethnologue, but focused on technology: each language and writing system gets its own page listing where it is spoken, how it is written, how endangered it is, and which identification models, corpora, language models and OCR systems support it today.
Projects
Eight open releases, one pipeline. Every model, dataset and tool is free to use.
GlotLID
Language identification with support for more than 2,000 labels.
Glot500
Scaling multilingual corpora and language models to 500+ languages.
GlotScript
A resource and tool for writing system identification.
GlotCC
An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.
GlotWeb
Web indexing for minority languages.
GlotOCR Bench
OCR models still struggle beyond a handful of Unicode scripts.
GlotStoryBook
Openly licensed storybooks for 180 languages.
GlotSparse
Text for sparsely resourced languages.
GlotSuite
GlotSuite is the family of research releases at the heart of CommonGlot. Each project solves one link in the chain from raw web text to working language technology, and each one feeds the next.
Is your language missing?
Corrections, data and evaluations from language communities are the fastest way to improve coverage. Open an issue or send a pull request — every booklet links to its sources.