Model · mpt-7b-chat
MPT-7B-Chat
Developer: MosaicML
Availability: removed · checked 2026-09-24 · source
HF API returned HTTP 401. Card read from an Internet Archive snapshot dated 2023-07-18.
Raw record: /data/models/mpt-7b-chat.json
Fields
- id
- mpt-7b-chat
- identifiers
- huggingface
- mosaicml/mpt-7b-chat
- developer
- MosaicML
- release_date
- 2023-05-05recorded · sourcenote: Release post (MosaicML/Databricks blog) and the archived card ('Model Date May 5, 2023'). HF API returned 401 for this repo (mosaicml/ org unreachable, same as mpt-7b).
- weights_status
- unknown
- availability
- removed · checked 2026-09-24 · sourcenote: HF API returned HTTP 401. Card read from an Internet Archive snapshot dated 2023-07-18.
- license
- CC-By-NC-SA-4.0recorded · sourcenote: Blog: 'CC-By-NC-SA-4.0 (non-commercial use only)'; the archived card says the same. Differs from mpt-7b's Apache-2.0.
- architecture
- family
- decoder_only
- n_layers
- 32recorded · sourcenote: Archived card architecture table: n_layers 32, n_heads 32, d_model 4096, vocab size 50432, sequence length 2048.
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 50432recorded · source
- context_length
- 2048recorded · source
- positional_encoding
- ALiBi (no position embeddings)recorded · sourcenote: Blog: MPT 'replaces positional embeddings with ALiBi'. Same as mpt-7b-instruct's record.
- note
- Archived card: 'follows a modified decoder-only transformer architecture'.
- training_data
- MPT-7B finetuned on the ShareGPT-Vicuna, HC3, Alpaca, HH-RLHF and Evol-Instruct datasets.recorded · sourcenote: Archived card: 'It was built by finetuning MPT-7B on the ShareGPT-Vicuna, HC3, Alpaca, HH-RLHF, and Evol-Instruct datasets.' Training run: 8.2 hours on 8 A100-80GB, then 6.7 hours on 32 A100-40GB.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (mosaicml/mpt-7b-chat@None) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date, license, architecture (from archived card), training data, 3 edges prepared; sources: archived card 2023-07-18, MosaicML/Databricks blog · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (3 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → MPT-7B declared source Blog: MPT-7B-Chat 'built by finetuning MPT-7B'; archived card: 'built by finetuning MPT-7B on ...'.
Training data
- trained_on → ShareGPT conversations (LMSYS collection) declared source Archived card names 'ShareGPT-Vicuna' (link jeffwan/sharegpt_vicuna, unreachable). The `sharegpt-vicuna` record is the LMSYS collection; mapping by name. Reviewer: confirm the id mapping.
- trained_on → Alpaca instruction data (52K) declared source Archived card lists Alpaca (link tatsu-lab/alpaca).
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.