Benchmark · arXiv 2026

GlotOCR Bench

OCR models still struggle beyond a handful of Unicode scripts.

Overview

GlotOCR Bench renders text in 150+ Unicode scripts with Google Fonts, in clean and degraded “old document” styles, and evaluates open and proprietary OCR models with CER, Acc@k and script accuracy. Results are broken down by script and by resource tier, revealing a steep drop beyond Latin and a few major scripts.

Features

  • Per-script and per-language results
  • High / mid / low resource tiers
  • Plain and old-document rendering profiles
  • Public leaderboard and released model outputs

Links