Corpus · NeurIPS 2024
GlotCC
An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.
Overview
GlotCC is a document-level corpus extracted from CommonCrawl with GlotLID and an open fork of the Ungoliant pipeline. The latest version covers more than 1,000 languages and applies quality filters adopted from C4, CCNet, MADLAD-400, RedPajama-v2, FineWeb, Dolma and others.
Features
- Document-level text for 1,000+ languages
- Reproducible open pipeline built on Ungoliant
- Language identification by GlotLID
- Filters adopted from major web corpora