Benchmark · arXiv 2026
GlotOCR Bench
OCR models still struggle beyond a handful of Unicode scripts.
Overview
GlotOCR Bench renders text in 150+ Unicode scripts with Google Fonts, in clean and degraded “old document” styles, and evaluates open and proprietary OCR models with CER, Acc@k and script accuracy. Results are broken down by script and by resource tier, revealing a steep drop beyond Latin and a few major scripts.
Features
- Per-script and per-language results
- High / mid / low resource tiers
- Plain and old-document rendering profiles
- Public leaderboard and released model outputs