Everything on this site is a JSON file.
The site is static and every page renders from the files below. Use them directly, mirror them, or rebuild them from the upstream sources with one script.
Last built 2026-10-04 · 7,567 languages · 175 scripts
Files
Language index
data/languages/index.json
One compact row per language: ISO 639-3, name, macroarea, family, endangerment, scripts, coordinates, speakers and GlotSuite coverage flags.
- i
- ISO 639-3
- n
- name
- m
- macroarea
- f
- family
- e
- endangerment level (1–6)
- s
- scripts (ISO 15924)
- t
- coverage flags
- y
- latitude
- x
- longitude
- p
- speakers
Language details
data/languages/details/{letter}.json
Booklet data sharded by the first letter of the ISO code: endonym, countries, Glottocode, Wikidata ID, documentation level, auxiliary scripts, GlotLID per-label metrics and Glot500 labels.
Download (a.json)Scripts
data/scripts.json
One record per ISO 15924 script: name, type, direction, Unicode ranges, specimen, languages using it, GlotLID label count and per-model OCR results.
DownloadOCR models
data/ocr_models.json
GlotOCR Bench summary per model: overall, high/mid/low-tier accuracy and macro CER.
DownloadProjects
data/projects.json
GlotSuite projects with descriptions, links, usage snippets and BibTeX.
DownloadVocabularies
data/taxonomy.json
Labels and colours for macroareas, endangerment levels, coverage flags and script types.
DownloadRebuild from source
git clone --depth 1 https://github.com/glottolog/glottolog-cldf src/glottolog-cldf
for r in GlotScript GlotLID GlotWeb GlotOCR-bench; do
git clone --depth 1 https://github.com/cisnlp/$r src/$r
done
pip install pycountry
python3 tools/build_data.py --sources srcSources & licences
| Source | Used for | Licence |
|---|---|---|
| Glottolog | Names, Glottocodes, families, macroareas, coordinates, endangerment, documentation level | CC BY 4.0 |
| GlotScript-R | Writing systems per language, Unicode ranges per script | MIT |
| GlotLID v3 | Supported labels and per-label F1 / precision / recall | Apache-2.0 |
| LinguaMeta (via GlotWeb) | Endonyms, speaker estimates, Wikidata IDs | CC BY 4.0 |
| Glot500 label list (via GlotWeb) | Languages covered by Glot500 | Apache-2.0 |
| GlotOCR Bench | Per-script OCR accuracy and CER for 14 models | see repository |