← Contents Appendix

Appendix

Model by Model

Every lecture in this series draws its numbers from somewhere. This appendix is that somewhere, collected into one table per laboratory: release date, parameter count, attention mechanism, context length, training precision, and — the two columns most reports leave out — what precision the weights were actually released in, and what a server needs to run them. A blank cell is not an omission on this page's part; it is the fact. The laboratory did not say. Lecture 23 argues that the shape of the gaps is itself the most honest thing a report tells you — this table is where that argument becomes checkable.

How to read a row

Total / Active is parameters at rest vs. parameters touched per token — the gap between the two is Volume VII's subject. Attention abbreviates the mechanism Volumes I–VI name in full (MHA, GQA, MLA, hybrid ratios, sparse variants). Train precision is the number format pretraining actually ran in, when a report says so. Released / inference is the format the public weights ship in, or the format a report specifically claims for serving — a different question from training precision, and the one most reports skip entirely. unconfirmed marks a figure from secondary coverage rather than a primary paper, model card, or official repository.


DeepSeek

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
V3Dec 2024671B / 37BMLA128KFP8 mixed-precisionnot disclosed2412.19437
R1Jan 2025671B / 37B (V3 backbone)MLA128Knot disclosednot disclosed2501.12948
V3.2 / V3.2-ExpDec 2025671B / 37BMLA + DSA128Knot disclosednot disclosed2512.02556
V4 Pro / FlashApr 20261.6T / 49B · 284B / 13Bhybrid sparse (unconfirmed detail)1Mnot disclosednot disclosedrelease note unconfirmed

Qwen — Alibaba

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Qwen3 (dense + MoE)May 20250.6B–32B dense · up to 235B / 22B MoEGQA128K (extended)not disclosednot disclosed2505.09388
Qwen3-NextSep 202580B / 3BGated DeltaNet : full, 3:1not disclosednot disclosednot disclosedblog only
Qwen3.5-397B-A17BFeb 2026397B / 17BGated DeltaNet : gated full, 3:1262Knot disclosednot disclosedmodel card — no technical report as of Feb 2026

Kimi — Moonshot AI

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
K2Jul 20251.04T / 32B (384 experts)MLA128Knot disclosednot disclosed2507.20534
K2 ThinkingNov 20251.04T / 32BMLA128Knot disclosedINT4, quantization-aware trainingmodel card
K2.5Jan 2026~1T / ~32BMLAnot disclosednot disclosednot disclosedmodel page
K316 Jul 20262.8T / ~16 of 896 experts (≈1.8%)Kimi Delta Attention (linear hybrid) + MLA1Mnot disclosednot disclosedtechnical report not yet published unconfirmed

Meta — Llama

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Llama 4 ScoutApr 2025109B / 17B (16 experts)iRoPE (RoPE local + NoPE global)10MFP8not disclosedblog only, no paper
Llama 4 MaverickApr 2025400B / 17B (128 experts)iRoPEnot disclosedFP8not disclosedblog only
Llama 4 Behemothnever shipped~2T / 288B (previewed)not disclosednot disclosednot disclosednot disclosedpreview only

Google — Gemini & Gemma

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Gemini 2.5 Pro / Flash2025sparse MoE, size not disclosednot disclosed1Mnot disclosednot disclosed2507.06261
Gemini 3 ProNov 2025sparse MoE, size not disclosednot disclosed1M in / 64K outnot disclosednot disclosedmodel card only
Gemini 3.1 ProFeb 2026sparse MoE, size not disclosednot disclosed1Mnot disclosednot disclosedmodel card only
Gemma 4Jun 2026 (report)31B dense · 26B / 4B MoEnot disclosed256K (31B)not disclosedApache 2.0 weights, format not disclosed2607.02770

OpenAI

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
GPT-5Aug 2025not disclosed (router of models)not disclosednot disclosednot disclosednot disclosedsystem card only
gpt-oss-120bAug 2025117B / 5.1Balternating dense + sliding window, learned sinks131K (YaRN-extended)not disclosedMXFP4 (expert weights) — fits one 80GB GPU2508.10925
gpt-oss-20bAug 202521B / 3.6Balternating dense + sliding window, learned sinks131Knot disclosedMXFP4 (expert weights)2508.10925
GPT-5.1 / 5.5 / 5.6 (Luna, Terra, Sol)Nov 2025 – Jul 2026not disclosednot disclosednot disclosednot disclosednot disclosedsystem cards / preview posts only

Anthropic

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Opus 4.1 / Sonnet 4.5 / Opus 4.5Aug–Nov 2025not disclosednot disclosednot disclosednot disclosednot disclosedsystem cards only
Opus 4.8May 2026not disclosednot disclosednot disclosednot disclosednot disclosedsystem card only
Fable 5 / Mythos 5Jun 2026not disclosednot disclosed1M in / 128K out (third-party reported)not disclosednot disclosedsystem card only

Zero architecture rows in this table have a single disclosed cell — that emptiness, across an entire laboratory, is itself the data point Lecture 23 asks you to notice.

Mistral

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
MagistralJun 2025(Mistral Medium 3 backbone)not disclosednot disclosednot disclosednot disclosed2506.10910
Mistral Large 3Dec 2025675B / 41Bnot disclosed256Knot disclosedApache 2.0 weights, format not disclosedblog + weights, no paper

GLM — Zhipu / Z.ai

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
GLM-4.5 (+ Air)Jul 2025355B / 32B · Air 106B / 12Bhybrid thinking modesnot disclosednot disclosednot disclosed2508.06471
GLM-5Feb 2026744B / 44BMLA + DSA (adopted from DeepSeek)200Knot disclosednot disclosed2602.15763

MiniMax

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
M1Jun 2025456B / 45.9BLightning (linear) hybrid1Mnot disclosednot disclosed2506.13585
M2model Oct 2025, report May 2026230B / 10B (256 experts)full attention (deliberate reversion)192Knot disclosednot disclosed2605.26494
M3Jun 2026428B / 22BGQA + MiniMax Sparse Attention1Mnot disclosednot disclosed2606.13392

xAI — Grok

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Grok 4Jul 2025~3T total (reported)not disclosednot disclosednot disclosednot disclosedmodel card only
Grok 4.1 / 4 FastNov 2025not disclosednot disclosed2M (Fast)not disclosednot disclosedmodel cards only
Grok 4.5Jul 2026~1.5T "V9 foundation" (reported)not disclosednot disclosednot disclosednot disclosedsingle social post unconfirmed

NVIDIA — Nemotron

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Nemotron 3 NanoDec 202530B / 3BMamba-2 hybrid1M-classNVFP4NVFP42512.20856
Nemotron 3 SuperDec 2025~120B / 12BMamba-2 hybrid1M-classNVFP4NVFP42512.20856
Nemotron 3 UltraDec 2025550B / 55BMamba-2 hybrid1M-classNVFP4 — first 4-bit pretrain at this scaleNVFP42512.20856

Ai2 — OLMo

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
OLMo 3 (7B, 32B)Nov 20257B / 7B · 32B / 32B (dense)not disclosednot discloseddisclosed in full training code (not summarised here)full weights + every checkpoint, native formatblog + full repo — the only fully reproducible release in this table

Other notable releases

ModelReleasedTotal / ActiveAttentionContextTrain precisionReleased / inferenceSource
Falcon-H1 (TII)2025not disclosedparallel Mamba-2 + attention, same blocknot disclosednot disclosednot disclosed2507.22448
Jamba 1.7 (AI21)2025not disclosedMamba + attention + MoE hybrid256Knot disclosednot disclosedoriginal Jamba: 2403.19887
Command A (Cohere)Mar 2025111B densenot disclosed256Knot disclosednot disclosed2504.00698
SmolLM3 (Hugging Face)Jul 20253B denseNoPE on every 4th layer128K (staged 4K→64K→128K)disclosed in full training codefully open weights + data + codeblog
Hunyuan 3.0 "Hy3" (Tencent)Jul 2026295B / ~21B (192 experts)not disclosed256Knot disclosedFP8 releasecoverage only unconfirmed
ERNIE 5.0 (Baidu)Feb 2026not disclosednot disclosednot disclosednot disclosednot disclosed2602.04705

What the blanks add up to

Count the disclosed "Released / inference" cells above: four labs state a released or serving precision in a primary source — OpenAI (MXFP4, gpt-oss only, its one forced-open release), NVIDIA (NVFP4, the whole Nemotron 3 line), Kimi (INT4 QAT, the Thinking variant only), and Ai2 (native, because OLMo 3 ships everything). Every frontier closed lab — OpenAI's flagship line, Anthropic, Google's Gemini, xAI — discloses nothing about the number format its weights actually run in. That is not this page failing to find the fact. It is the fact.

Where these rows come from

  • Primary sources are linked per row above; anything without a direct link was verified against the model's official blog, card, or repository rather than a paper.
  • Volume VII (lectures 15–16) explains the architecture column in full; Volume VIII (17–18) explains the two precision columns; Volume XI (23–24) is the argument this table exists to support.
  • Sebastian Raschka's LLM Architecture Gallery — the closest existing equivalent to this table, and worth cross-checking against.