ROOTS corpus

E1312454 UNEXPLORED

The ROOTS corpus is a massive, multilingual text dataset compiled to train and evaluate large language models such as those in the BLOOM project.

All labels observed (1)

Label Occurrences
ROOTS corpus canonical 1

How this entity was disambiguated

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Bloom trainingDataSource ROOTS corpus