Model · cerebras-gpt-13b
Cerebras-GPT 13B
Developer: Cerebras Systems (Nolan Dey, Gurpreet Gosal, Zhiming (Charles) Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness; trained on the Cerebras Wafer-Scale Cluster)
Availability: removed · checked 2026-09-24 · source
HF API returned HTTP 401.
Raw record: /data/models/cerebras-gpt-13b.json
Fields
- id
- cerebras-gpt-13b
- identifiers
- huggingface
- cerebras/Cerebras-GPT-13B
- developer
- Cerebras Systems (Nolan Dey, Gurpreet Gosal, Zhiming (Charles) Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness; trained on the Cerebras Wafer-Scale Cluster)
- release_date
- 2023-03-28recorded · sourcenote: Cerebras blog post dated Mar 28 2023; paper arXiv 2304.03208 v1 is dated 6 Apr 2023.
- weights_status
- unknown
- availability
- removed · checked 2026-09-24 · sourcenote: HF API returned HTTP 401.
- license
- Apache-2.0recorded · sourcenote: Blog: 'All models, weights, and checkpoints are available on Hugging Face and GitHub under the Apache 2.0 license.' HF repo unreachable (HTTP 401) so the card was not read.
- architecture
- family
- decoder_only
- note
- HF config not retrieved (repo unreachable: HTTP 401). Values from paper Table 1 (13B row: d_model 5120, n_layers 40, d_head 128, d_ffn 20480) and section 2.1. n_heads is derived (5120 / 128 = 40); the paper gives d_head, not the head count. Positional encoding is not stated in the sections read.
- n_layers
- 40recorded · sourcenote: Paper Table 1.
- hidden_size
- 5120recorded · sourcenote: Paper Table 1 d_model.
- n_heads
- 40recorded · sourcenote: Derived: d_model 5120 / d_head 128; paper Table 1 does not list the count directly.
- vocab_size
- 50257recorded · sourcenote: Paper sec. 2.2: GPT-2 BPE vocabulary of size 50257.
- context_length
- 2048recorded · sourcenote: Paper sec. 2.1: 'maximum sequence length of 2048 tokens'.
- positional_encoding
- nullnot_recorded
- training_data
- The Pile (Gao et al., 2020), no deduplication; 257.1B tokens (about 20 tokens per parameter, Chinchilla-optimal).recorded · sourcenote: Paper: 'We train Cerebras-GPT models on the Eleuther Pile dataset following DeepMind Chinchilla scaling rules'; 'We do not perform deduplication of Pile'; Table 1 total tokens 257.1B for 13B.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (cerebras/Cerebras-GPT-13B@None) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date (blog), license (blog), architecture (paper), training data, 2 edges (trained_on the-pile; design_follows gpt-3); sources: arXiv 2304.03208, Cerebras blog · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Training data
- trained_on → The Pile declared source Paper abstract and sec. 2.2: trained on the Pile with the provided train/test/validation splits; no deduplication.
Design
- design_follows → GPT-3 (175B, 2020 paper) declared source Paper: 'Cerebras-GPT models have a GPT-3-like architecture, an autoregressive transformer decoder model. The main difference is that unlike GPT-3, which uses alternating dense and sparse-banded attention, we use dense attention in all decoder blocks.' 'GPT-3-like' with a stated exception -> design_follows (not same_architecture_retrained: the developer names a difference and does not say 'same').
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.