Open infrastructure · 7,000+ languages · 170+ scripts

Language tech for the long tail.

Most language technology serves a few dozen languages. CommonGlot builds and documents the open tools, corpora and benchmarks that reach the rest — and keeps a public record of what exists for every language and every script.

The directory in numbers

7,567languages in the directory
175writing systems tracked
2,102GlotLID language labels
157scripts with OCR benchmarks
242language families
Our mission

Every language deserves working technology.

We build and maintain free, open language technology for the thousands of languages that commercial systems ignore — and we publish an honest, inspectable record of what does and doesn't exist for each of them.

Read the mission →
◎

Identify

Know which language and script a text is in — the first step for any corpus.

⌘

Collect

Find and clean web text for languages that big crawls miss.

◆

Model

Train and release multilingual models that include the long tail.

▲

Evaluate

Benchmark systems per language and per script so gaps are visible.

What we deliver

Four kinds of open deliverables

Identification

Language & script identification

Models and libraries that tell you what language and writing system a piece of text is in — for 2,000+ language labels and every Unicode script.

  • fastText LID model (v1–v3)
  • Python script detector on PyPI
  • Language→script resource for 8,000+ languages
The directory

A booklet for every language and script

Like Glottolog or Ethnologue, but focused on technology: each language and writing system gets its own page listing where it is spoken, how it is written, how endangered it is, and which identification models, corpora, language models and OCR systems support it today.

GlotSuite

GlotSuite

GlotSuite is the family of research releases at the heart of CommonGlot. Each project solves one link in the chain from raw web text to working language technology, and each one feeds the next.

Is your language missing?

Corrections, data and evaluations from language communities are the fastest way to improve coverage. Open an issue or send a pull request — every booklet links to its sources.

How to contribute