Dataset · orca-1-data
Orca 1 explanation-tuning data (FLAN-5M / FLAN-1M)
Builder: Microsoft Research (Orca 1 authors)
Availability: unknown · checked 2026-09-24 · source
Neither paper says the training data was released: Orca 1 announces only a weight diff ('publicly release a diff of the model weights'), Orca 2 says 'We make Orca 2 weights publicly available'. A search of the Hugging Face microsoft org for 'orca' on 2026-09-24 returned only microsoft/orca-math-word-problems-200k and microsoft/orca-agentinstruct-1M-v1, neither of which is described as this data. Absence from a search is not proof of non-release, so recorded unknown, not never_released. OpenOrca (id openorca) is a third-party reproduction of Orca 1's recipe, not this data.
Raw record: /data/datasets/orca-1-data.json
Fields
- id
- orca-1-data
- builder
- Microsoft Research (Orca 1 authors)
- release_date
- 2023-06-05partial · sourcenote: arXiv v1 date of the Orca paper; the data itself has no separate release date.
- availability
- unknown · checked 2026-09-24 · sourcenote: Neither paper says the training data was released: Orca 1 announces only a weight diff ('publicly release a diff of the model weights'), Orca 2 says 'We make Orca 2 weights publicly available'. A search of the Hugging Face microsoft org for 'orca' on 2026-09-24 returned only microsoft/orca-math-word-problems-200k and microsoft/orca-agentinstruct-1M-v1, neither of which is described as this data. Absence from a search is not proof of non-release, so recorded unknown, not never_released. OpenOrca (id openorca) is a third-party reproduction of Orca 1's recipe, not this data.
- content
- 'We generate 5 million instructions (queries augmented with system messages) referred as FLAN-5M ... We further randomly sample 1 million queries from FLAN-5M to create another split, referred as FLAN-1M. We use Azure OpenAI API to collect ChatGPT (GPT-3.5-turbo) responses to FLAN-5M, and GPT-4 responses to FLAN-1M.' Orca 2's paper calls these '5 million ChatGPT data from Orca 1' and '1 million GPT-4 data from Orca 1'.recorded · sourcenote: Prompts are drawn from the FLAN-v2 collection (Orca 1 paper, section on scaling tasks). Two response sets, two teachers: ChatGPT wrote the 5M responses, GPT-4 the 1M subset. The ChatGPT snapshot is not_recorded.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (ruling: 3-tier panel (2/3 dataset routing), see session log) · by:wilson-pruitt + claude (Sonnet 5) ·
Parents
Influence without weights
- distilled_from_outputs → ChatGPT (Nov. 2022 launch model) declared source Builder's own paper: 'We use Azure OpenAI API to collect ChatGPT (GPT-3.5-turbo) responses to FLAN-5M'. Snapshot not_recorded.
- distilled_from_outputs → GPT-4 (Mar. 2023) declared source Builder's own paper: '... and GPT-4 responses to FLAN-1M' (1M queries sampled from the 5M).
Children
Training data
- ← trained_on Orca 2 7B declared source Orca 2 paper 4.2: 'We then train on 5 million ChatGPT data from Orca 1 for 3 epochs. Then we train on the combination of 1 million GPT-4 data from Orca 1 and Orca 2's 817K data for 4 epochs.'
Read in
No station on the reading path has touched this record yet.