Who we are
The CommonGlot Foundation is an open, non-commercial initiative that grew out of the GlotSuite research line at the Center for Information and Language Processing (CIS), LMU Munich, in collaboration with Sorbonne Université and CNRS. Like Common Crawl does for web data, it aims to keep language technology for every language open and freely available.
People behind GlotSuite
Amir Hossein Kargaran
Lead of GlotLID, GlotScript, GlotCC, GlotOCR Bench · LMU Munich
Hinrich Schütze
Co-author across GlotSuite · LMU Munich
François Yvon
Co-author across GlotSuite · Sorbonne Université, CNRS
Ayyoob Imani
Lead of Glot500, co-author of GlotLID · LMU Munich
Nafiseh Nikeghbal
Co-author of GlotOCR Bench
Abdullah Al Sefat
Lead of GlotWeb
Plus the many co-authors of Glot500 and the open-source contributors who report issues and send data.
How to contribute
Report an error
Wrong name, script or coverage flag? Open an issue on the site repository and link the booklet.
Share data
Point us to openly licensed text in your language, or to websites GlotWeb should index.
Evaluate
Test GlotLID, Glot500 or OCR models on your language and share the results.
Improve the site
Everything is static HTML + JSON. Pull requests are welcome.
Questions
What is the CommonGlot Foundation?
An open, non-commercial initiative — inspired by Common Crawl — that develops and maintains free language technology and data for the world's under-served languages.
Where does the directory data come from?
From Glottolog, GlotScript, LinguaMeta and GlotSuite releases, combined by a public build script. See the Open data page for every source and licence.
A language says “not covered” but I know a model exists.
The directory currently tracks GlotSuite coverage plus OCR benchmark results. Please open an issue — adding other open technologies is on the roadmap.
How do I cite this?
Each project page has a BibTeX entry with a copy button. Please cite the specific project you use.