WebText dataset

E99319

The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.

AI illustration

How this image was made

AI-generated illustration of WebText dataset

This AI-generated illustration was produced by black-forest-labs/FLUX.2-dev (1024x1024) from a prompt written by openai/gpt-oss-120b from the entity's label + description.

Prompt

Generate an image of the WebText dataset (The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.)

All labels observed (2)

Label Occurrences
OpenWebText 1
WebText dataset canonical 1

How this entity was disambiguated

Statements (49)

Predicate Object
instanceOf language model training dataset ⓘ
text corpus ⓘ
access not fully open to public download ⓘ
associatedWith GPT-2 ⓘ
collectionMethod crawling URLs extracted from Reddit ⓘ
comparedWith Wikipedia-only training corpora ⓘ
contains Wikipedia pages ⓘ
articles ⓘ
code snippets ⓘ
dialogue-like text ⓘ
news ⓘ
online books ⓘ
stories ⓘ
technical documentation ⓘ
web documents ⓘ
web forum discussions ⓘ
curatedBy OpenAI researchers ⓘ
curationFocus high-quality internet text ⓘ
dataModality natural language text ⓘ
dataSource outbound links from Reddit ⓘ
web pages ⓘ
developer OpenAI ⓘ
domain web text ⓘ
excludes low-quality spam pages ⓘ
non-text-heavy pages ⓘ
goal capture broad distribution of internet text ⓘ
improve generalization of language models ⓘ
influenced later web-scale language modeling datasets ⓘ
language English ⓘ
license not publicly released as a full dataset ⓘ
organization OpenAI ⓘ
preprocessingStep deduplication of documents ⓘ
filtering low-quality pages ⓘ
tokenization ⓘ
publication Language Models are Unsupervised Multitask Learners ⓘ
publicationYear 2019 ⓘ
relatedTo OpenAI GPT models ⓘ
linked to: OpenAI API platform
releasedBy OpenAI ⓘ
scale large-scale ⓘ
selectionCriterion filtering for high-quality content ⓘ
links from Reddit with high karma ⓘ
sizeDescription on the order of billions of tokens ⓘ
topicCoverage diverse internet topics ⓘ
trainingObjective next-token prediction ⓘ
usedFor training GPT-2 ⓘ
training large language models ⓘ
unsupervised language modeling ⓘ
usedIn evaluation of GPT-2 capabilities ⓘ
research on zero-shot learning with language models ⓘ

How these facts were elicited

Referenced by (2)

Full triples — surface form annotated when it differs from this entity's canonical label.

GPT-2 → trainingDataSource → WebText dataset ⓘ
RoBERTa → trainingDataSource → OpenWebText ⓘ
linked to: WebText dataset