Dataset · ultrafeedback
UltraFeedback
Builder: OpenBMB / Tsinghua University
Availability: available · checked 2026-09-24 · source
Raw record: /data/datasets/ultrafeedback.json
Fields
- id
- ultrafeedback
- identifiers
- huggingface
- openbmb/UltraFeedback
- builder
- OpenBMB / Tsinghua University
- release_date
- 2023-10-02partial · sourcenote: arXiv v1 date of the paper.
- availability
- available · checked 2026-09-24 · source
- content
- Preference data: instructions answered by a pool of models, with GPT-4 employed 'to offer detailed feedback in both numerical and textual forms.'recorded · sourcenote: The answering models are not recorded here; only GPT-4's role as judge is.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (dataset records ruling) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- feedback_from → GPT-4 (Mar. 2023) declared source Builder's own paper: GPT-4 employed 'to offer detailed feedback in both numerical and textual forms.'
Children
Training data
- ← trained_on Zephyr 7B β declared source DPO step on GPT-4's rankings of model completions.
- ← trained_on zephyr-7b-alpha declared_by_uploader source HF card metadata datasets: openbmb/UltraFeedback. Often incomplete; confirm which training stage used it.
- ← trained_on Tulu 2 DPO 7B declared source Card metadata `datasets: HuggingFaceH4/ultrafeedback_binarized` (a filtered, binarized UltraFeedback); paper: DPO on 'a filtered and binarized form of UltraFeedback'. Recorded like Zephyr: trained_on the UltraFeedback record, which itself carries feedback_from gpt-4. Tulu's own paper notes UltraFeedback used TruthfulQA prompts (contamination caveat for evaluation).
Read in
No station on the reading path has touched this record yet.