Colossal Clean Crawled Corpus

E1312424 UNEXPLORED

The Colossal Clean Crawled Corpus (C4) is a massive, cleaned web-text dataset widely used to train large language models and other state-of-the-art NLP systems.

All labels observed (1)

Label Occurrences
Colossal Clean Crawled Corpus canonical 1

How this entity was disambiguated

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

T5 trainingData Colossal Clean Crawled Corpus