Projects
Eight open releases, one pipeline. Every model, dataset and tool is free to use.
All projects
GlotLID
Language identification with support for more than 2,000 labels.
Glot500
Scaling multilingual corpora and language models to 500+ languages.
GlotScript
A resource and tool for writing system identification.
GlotCC
An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.
GlotWeb
Web indexing for minority languages.
GlotOCR Bench
OCR models still struggle beyond a handful of Unicode scripts.
GlotStoryBook
Openly licensed storybooks for 180 languages.
GlotSparse
Text for sparsely resourced languages.
GlotCC
An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.
GlotStoryBook
Openly licensed storybooks for 180 languages.
GlotSparse
Text for sparsely resourced languages.
GlotLID labels the language, GlotScript the writing system.
GlotCC and GlotWeb gather and filter web text with those labels.
GlotCC
An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.
GlotWeb
Web indexing for minority languages.
GlotStoryBook
Openly licensed storybooks for 180 languages.
GlotSparse
Text for sparsely resourced languages.
Glot500 learns from the collected corpora.
GlotOCR Bench and per-language results show where gaps remain.