Dataset · baize-sdf
Baize SDF data (ChatGPT-ranked self-generations)
Builder: UC San Diego / Sun Yat-sen University (Baize authors)
Availability: unknown · checked 2026-09-24 · source
Repo returned HTTP 200. The paper says 'we are also releasing the fine-tuning corpus' but does not say whether the SDF set is in it; not confirmed.
Raw record: /data/datasets/baize-sdf.json
Fields
- id
- baize-sdf
- identifiers
- builder
- UC San Diego / Sun Yat-sen University (Baize authors)
- release_date
- 2023-04-03partial · sourcenote: arXiv v1 date of the Baize paper, which introduces SDF; the release date of any SDF data is not recorded.
- availability
- unknown · checked 2026-09-24 · sourcenote: Repo returned HTTP 200. The paper says 'we are also releasing the fine-tuning corpus' but does not say whether the SDF set is in it; not confirmed.
- content
- Paper, Self-Distillation with Feedback (SDF): 'we use the resulted Baize v1.5 models to generate four responses for each instruction from the Quora dataset mentioned in Table 2. We then engage ChatGPT using the prompt provided in Appendix C to rank generate responses for self-distillation. Finally, we select the best response ranked by ChatGPT to finetune the model.' The training targets are the model's own generations (Baize v1.5); ChatGPT supplied only the ranking.recorded · sourcenote: Quora instructions only, per the paper's SDF passage; size not recorded.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (ruling: Wilson, after 3-tier panel split, see session log) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- feedback_from → ChatGPT (Nov. 2022 launch model) declared source Paper: 'we use the resulted Baize v1.5 models to generate four responses for each instruction from the Quora dataset mentioned in Table 2. We then engage ChatGPT using the prompt provided in Appendix C to rank generate responses for self-distillation. Finally, we select the best response ranked by ChatGPT to finetune the model.' ChatGPT ranked candidates; the selected text is Baize v1.5's own output, so this is feedback (judgments), not distilled_from_outputs. ChatGPT snapshot not_recorded.
Children
Training data
- ← trained_on Baize v2 7B declared source Card: v2 trained with 'self-distillation with feedback (SDF)'. Paper Table 3 lists Baize-v2-7B as SDF applied to Baize-v1.5-7B: 'we apply new LoRA modules to all linear layers in Baize v1.5'. SDF data is Baize v1.5's own generations for Quora instructions, ranked by ChatGPT.
Read in
No station on the reading path has touched this record yet.