Dataset · refinedweb
RefinedWeb
Builder: Technology Innovation Institute (TII)
Availability: partial · checked 2026-09-24 · source
Paper: 'We publicly release an extract of 600 billion tokens' of the five-trillion-token dataset.
Raw record: /data/datasets/refinedweb.json
Fields
- id
- refinedweb
- identifiers
- huggingface
- tiiuae/falcon-refinedweb
- builder
- Technology Innovation Institute (TII)
- release_date
- 2023-06-01partial · sourcenote: arXiv v1 date of the paper.
- availability
- partial · checked 2026-09-24 · sourcenote: Paper: 'We publicly release an extract of 600 billion tokens' of the five-trillion-token dataset.
- content
- Filtered, deduplicated English web text from CommonCrawl, ~5 trillion tokens.recorded · source
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (dataset records ruling) · by:wilson-pruitt + claude ·
Parents
No edges recorded.
Children
Training data
- ← trained_on Falcon 7B declared source RefinedWeb-English is 79% of 1,500B training tokens; the rest is curated corpora (not recorded as datasets).
- ← trained_on falcon-40b declared_by_uploader source HF card metadata datasets: tiiuae/falcon-refinedweb. Often incomplete; confirm which training stage used it.
- ← trained_on falcon-7b-instruct declared_by_uploader source Card: RefinedWeb-English is 5% (13M of 250M) of the fine-tuning mix.
- ← trained_on phi-2 declared source Card: training mix includes 'filtered web data from Falcon RefinedWeb and SlimPajama, which was assessed by AOAI GPT-4'.
- ← trained_on falcon-40b-instruct declared_by_uploader source Card: '150M tokens from Baize mixed with 5% of RefinedWeb data.'
Read in
No station on the reading path has touched this record yet.