Corpus · NeurIPS 2024

GlotCC

An open, broad-coverage CommonCrawl corpus and pipeline for minority languages.

Overview

GlotCC is a document-level corpus extracted from CommonCrawl with GlotLID and an open fork of the Ungoliant pipeline. The latest version covers more than 1,000 languages and applies quality filters adopted from C4, CCNet, MADLAD-400, RedPajama-v2, FineWeb, Dolma and others.

Features

  • Document-level text for 1,000+ languages
  • Reproducible open pipeline built on Ungoliant
  • Language identification by GlotLID
  • Filters adopted from major web corpora

Links