Model · gpt-jt-6b-v1
GPT-JT-6B-v1
Developer: Together Computer
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/gpt-jt-6b-v1.json
Fields
- id
- gpt-jt-6b-v1
- identifiers
- huggingface
- togethercomputer/GPT-JT-6B-v1
- developer
- Together Computer
- release_date
- 2022-11-29recorded · sourcenote: Together release post, 'Published 11/29/2022'. The HF repo's first commit is earlier (2022-11-24, HF API); the post is the developer's release announcement.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- apache-2.0recorded · sourcenote: Card: 'License: Apache License 2.0'. Uploader is the developer.
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['GPTJForCausalLM'].
- n_layers
- 28recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 16recorded · source
- vocab_size
- 50400recorded · source
- context_length
- 2048recorded · source
- positional_encoding
- rotary (RoPE), partial: 64 of 256 dims per headrecorded · sourcenote: Propagated from gpt-j-6b (GPT-JT is 'a fork of' GPT-J with the same config). The UL2 prefix-mask objective changes the attention mask used in training, not the position encoding.
- training_data
- 3.53B fine-tuning tokens in two stages: 2.62B tokens with the UL2 loss on the Pile, then 0.92B tokens mixing 5% Chain-of-Thought, 20% P3, 20% Natural Instructions and 55% the Pile.recorded · source
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (togethercomputer/GPT-JT-6B-v1@f34aa35f906895602c1f86f5685e598afdea8051) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date (blog), license, positional encoding (propagated from gpt-j-6b), training data, 2 edges accepted + 3 rejected; sources: card, Together blog · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → GPT-J 6B declared source Blog: 'A fork of GPT-J-6B, fine-tuned on 3.53 billion tokens'; card: 'a fork of EleutherAI's GPT-J (6B)'.
Training data
- trained_on → The Pile declared source Card: 2.62B tokens with UL2 loss on the Pile, and 55% of the second-stage mix.
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.