Stemma Machinarum

Model · gpt-neox-20b

GPT-NeoX-20B

Developer: EleutherAI
Availability: available · checked 2026-09-24 · source

Raw record: /data/models/gpt-neox-20b.json

Fields

id
gpt-neox-20b
identifiers
huggingface
EleutherAI/gpt-neox-20b
developer
EleutherAI
release_date
2022-04-14partial · source
note:  arXiv v1 date of the paper. EleutherAI announced the weights earlier; that date was not verified in-session.
weights_status
open
availability
available · checked 2026-09-24 · source
license
Apache 2.0recorded · source
architecture
family
decoder_only
note
Paper sec. 2.1: 'largely follows that of GPT-3' and 'almost identical to that of GPT-J'. Deviations: rotary on the first 25% of dims, attention and feed-forward computed in parallel, all-dense layers, new tokenizer trained on the Pile.
n_layers
44recorded · source
hidden_size
6144recorded · source
n_heads
64recorded · source
vocab_size
50432recorded · source
note:  Paper describes a new BPE tokenizer 'trained on the Pile' with vocabulary 50257; config's 50432 is the padded embedding size.
positional_encoding
rotary (RoPE), partial: first 25% of dimsrecorded · source
context_length
2048recorded · source
training_data
The Pile (EleutherAI's ~800GB curated English corpus)recorded · source
techniques
transformer-decoder
rotary-position-embedding
parallel-attention-ffn
primary_sources
https://arxiv.org/abs/2204.06745
https://huggingface.co/EleutherAI/gpt-neox-20b/blob/c292233c833e336628618a88a648727eb3dff0a7/README.md
https://huggingface.co/EleutherAI/gpt-neox-20b/raw/c292233c833e336628618a88a648727eb3dff0a7/config.json
record_history
date:2026-09-24 · change:created from primary sources (Phase 1 seed, batch 2) · by:wilson-pruitt + claude ·
date:2026-09-24 · change:availability checked and recorded · by:wilson-pruitt + claude ·

Parents

Training data

Design

Children

No edges recorded.

Read in

No station on the reading path has touched this record yet.