Model · openchat-3-5
OpenChat 3.5
Developer: OpenChat team (paper: Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, Yang Liu; Tsinghua University, Shanghai AI Laboratory, 01.AI); HF uploader imone
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/openchat-3-5.json
Fields
- id
- openchat-3-5
- identifiers
- huggingface
- openchat/openchat_3.5
- developer
- OpenChat team (paper: Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, Yang Liu; Tsinghua University, Shanghai AI Laboratory, 01.AI); HF uploader imone
- release_date
- 2023-11-01recorded · sourcenote: OpenChat GitHub README changelog: '[2023/11/01] We released the OpenChat-3.5-7B model'. HF repo created 2023-10-30.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Apache-2.0recorded · sourcenote: Card text: 'Our OpenChat 3.5 code and models are distributed under the Apache License 2.0.' (also card metadata). License file text not separately read.
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['MistralForCausalLM'].
- n_layers
- 32recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32002recorded · sourcenote: 32,002 vs 32,000 for Mistral-7B-v0.1: the config's `_name_or_path` is imone/Mistral_7B_with_EOT_token, a Mistral 7B variant with an added end-of-turn token.
- context_length
- 8192recorded · source
- n_kv_heads
- 8recorded · source
- positional_encoding
- rotary (RoPE)recorded · sourcenote: Architecture unchanged from mistral-7b-v0-1 (fine-tune, not a structural change).
- training_data
- A collection of publicly available instruction data with a custom processing pipeline, trained with C-RLFT (class-conditioned data sources, no preference labels). Notable subsets named on the card: OpenChat ShareGPT, OpenOrca with FLAN answers, Capybara (Pure-Dove, Verified-Camel, LessWrong-Amplify-Instruct), GOAT, Glaive, MetaMathQA, MathInstruct, OpenAssistant top-1.partial · sourcenote: Card 'Dataset Details': 'trained with C-RLFT on a collection of publicly available high-quality instruction data ... notable subsets included here'. The list is not stated to be complete.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (openchat/openchat_3.5@0fc98e324280bc4bf5d2c30ecf7b97b84fb8a19b) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date (README), license, training data (partial), positional encoding (propagated), 5 edge entries (fine_tuned_from mistral, trained_on openorca; rest rejected); sources: card, config.json, arXiv 2309.11235, OpenChat README · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → Mistral 7B (v0.1) declared_by_uploader source No card sentence names the base. Evidence: config.json `_name_or_path: imone/Mistral_7B_with_EOT_token`, `model_type: mistral`, dims identical to Mistral 7B, card tag `mistral`; and the Starling-LM-7B-alpha card (a third party) says 'Openchat 3.5 (based on Mistral-7B-v0.1)'. Uploader imone is the OpenChat author, so this is the developer's own HF metadata, not a prose statement. The intermediate 'Mistral_7B_with_EOT_token' checkpoint (vocab +2) was not inspected.
Training data
- trained_on → OpenOrca declared source Card lists imone/OpenOrca_FLAN; that dataset's card: 'the OpenOrca GPT4 subset with the original FLAN answers. Each even row contains the OpenOrca GPT4 answer, while each odd row contains the corresponding FLAN answer.' Mapped to the `openorca` record as a derived subset (a partial mapping; the GPT-4 half is the OpenOrca distillation of GPT-4).
Children
Weights descend
- ← fine_tuned_from Starling-LM-7B-alpha declared source Card: 'Finetuned from model: Openchat 3.5 (based on Mistral-7B-v0.1)'; 'a language model trained from Openchat 3.5 with reward model ... and policy optimization method APA'. Developers' own card. PARENT openchat-3-5 IS STAGED (promote it first).
Read in
No station on the reading path has touched this record yet.