Stemma Machinarum

Dataset · webtext

WebText

Builder: OpenAI
Availability: partial · checked 2026-09-24 · source
OpenAI released 250K documents from the WebText test set; the training set was not released.

Raw record: /data/datasets/webtext.json

Fields

id
webtext
builder
OpenAI
release_date
nullnot_recorded
availability
partial · checked 2026-09-24 · source
note:  OpenAI released 250K documents from the WebText test set; the training set was not released.
content
Millions of web pages linked from Reddit; the GPT-2 training corpus.recorded · source
note:  The gpt2-xl record carries the paper's fuller description (~8M documents, outbound links with karma >= 3).
primary_sources
https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
https://github.com/openai/gpt-2-output-dataset
record_history
date:2026-09-24 · change:created from primary sources (dataset records ruling) · by:wilson-pruitt + claude ·

Parents

No edges recorded.

Children

Training data

Read in