Our mission

Every language deserves working technology.

We build and maintain free, open language technology for the thousands of languages that commercial systems ignore — and we publish an honest, inspectable record of what does and doesn't exist for each of them.

Five pillars

◎

Identify

Know which language and script a text is in — the first step for any corpus.

⌘

Collect

Find and clean web text for languages that big crawls miss.

◆

Model

Train and release multilingual models that include the long tail.

▲

Evaluate

Benchmark systems per language and per script so gaps are visible.

▤

Document

Keep a public directory of what technology exists for each language.

Over 7,000 languages are spoken today, yet the web corpora, language models and OCR engines that power modern NLP cover only a small fraction of them. When a language can't be identified, it can't be collected; when it can't be collected, no model learns it; when no model learns it, its speakers are left out.

CommonGlot breaks that chain at each link. GlotLID and GlotScript identify languages and writing systems. GlotCC, GlotWeb and our smaller corpora collect text. Glot500 turns it into models. GlotOCR Bench measures how well today's systems read the world's scripts. Everything is released under open licences.

Alongside the tools, we maintain a directory that is part catalogue and part audit: one booklet per language and per script, built from public sources, showing exactly which technologies support it.

What we deliver

Four kinds of open deliverables

Identification

Language & script identification

Models and libraries that tell you what language and writing system a piece of text is in — for 2,000+ language labels and every Unicode script.

  • fastText LID model (v1–v3)
  • Python script detector on PyPI
  • Language→script resource for 8,000+ languages

Principles

01

Open by default

Code, models and data packaging are released under permissive licences (Apache-2.0, MIT, CC0) wherever the source allows.

02

Every claim has a source

Directory entries are generated from Glottolog, GlotScript, LinguaMeta and our own releases, and link back to them.

03

Coverage over headlines

We report per-language and per-script results, not just averages that hide the long tail.

04

Communities first

Speakers know their languages best. Corrections and contributions from communities take priority.