Volume IV · The Reality of Scale
Gigawatts, Not Just GPUs
Lecture 9 showed that useful compute is a fraction of peak hardware compute, and that fraction is disclosed with reasonable rigor by the labs that publish it. This lecture asks about the constraint sitting one level further out — electrical power and capital — and the honest answer is that this number is disclosed far less rigorously than anything in Lecture 1's chip spec sheets. That gap in disclosure quality is itself the lesson, not a footnote to it.
The mental model
Power and money are now load-bearing constraints on frontier AI, the same way FLOPS and memory bandwidth are — but unlike FLOPS and bandwidth, the power and cost numbers circulating about frontier training runs are mostly press-aggregated estimates layered on top of a much smaller set of primary disclosures. Knowing which is which is the actual skill this lecture teaches.
Stargate — what OpenAI itself has disclosed
Stargate is a joint OpenAI/Oracle/SoftBank infrastructure program, and the primary sources for it are OpenAI's own blog posts, not third-party reporting. OpenAI announced Stargate on January 21, 2025 with a stated target of $500B and 10GW of U.S. AI infrastructure buildout — that $500B/10GW figure is OpenAI's own number, describing a target, not a quantity already built. Separately, OpenAI has disclosed an Oracle/OpenAI agreement for an additional 4.5GW of Stargate capacity. Treat that 4.5GW figure as a distinct deal from the 10GW target, not a subset or update of it — the two numbers describe different agreements and should not be added together as if summing one running total.
Where press aggregation takes over from disclosure
Reporting about Stargate frequently states an aggregate running total — something like "~7GW currently planned or built" — and frequently follows a gigawatt figure with a human-scale comparison, such as "10GW is enough to power 8–10 million U.S. homes." Neither of those framings is something OpenAI itself states. Both are press-aggregated: a journalist or analyst summing OpenAI's individual deal announcements into a running total, or converting a gigawatt figure into a homes-powered comparison using a separate, unstated set of assumptions about average household consumption. State this plainly when citing either kind of figure — OpenAI's own per-deal numbers (the $500B/10GW target, the 4.5GW Oracle agreement) are primary; any aggregate total or homes-powered comparison built on top of them is press-sourced framing, not an OpenAI-stated fact.
DeepSeek-V3's cost, restated as the money side of scale
Lecture 9 already cited DeepSeek-V3's own disclosed training cost (arXiv:2412.19437): 2.788M H800 GPU-hours for full pretraining, with the paper's own cost estimate assuming $2/GPU-hour rental. That figure belongs here too, restated briefly, because it is the cleanest example in this lecture of a primary, paper-disclosed cost number — one lab stating its own training bill in its own paper, with the rate assumption stated alongside it, rather than a third party reconstructing a cost estimate from public cloud pricing.
xAI Colossus — flagged, not verified
Reporting on xAI's Colossus cluster commonly cites a power draw around 250MW. That number should be treated explicitly as press- and vendor-sourced — it traces to a Supermicro case study and trade-press coverage, not to an xAI engineering blog or formal disclosure with xAI's own name on the figure. Nothing in this lecture treats the 250MW number as verified; it is presented only as an example of the kind of figure that circulates confidently without a primary source behind it.
| Claim | Source | Status |
|---|---|---|
| Stargate: $500B / 10GW U.S. buildout target | OpenAI's own blog (Jan 21, 2025) | Primary |
| Stargate: +4.5GW Oracle/OpenAI capacity | OpenAI's own blog | Primary — a separate deal, don't sum with the 10GW target |
| "~7GW currently planned/built"; "10GW = 8–10M homes" | Press aggregation of the above | Press-sourced, not OpenAI-stated |
| DeepSeek-V3: 2.788M H800 GPU-hours, $2/GPU-hour | DeepSeek's own paper (arXiv:2412.19437) | Primary |
| xAI Colossus: ~250MW power draw | Supermicro case study + trade press | Press/vendor-sourced, unconfirmed by xAI |
The general pattern, stated directly
At time of writing, no major AI lab — OpenAI's own specific Stargate deal figures aside — has published a formal, self-authored disclosure of a specific frontier training run's power draw. Not Meta, not Google, not xAI. The discipline this lecture is teaching is narrow and mechanical: prefer citing only numbers a company put in its own blog post, SEC filing, or formal announcement, and flag everything else — which, for most cluster GPU-count and power-draw claims circulating in press coverage, means nearly all of them — as an estimate, not a verified fact.
Why this closes Part One
Lectures 1 through 9 described a machine whose theoretical limits (Lecture 1) and realized efficiency (Lecture 9) are both disclosed with real rigor — spec sheets, peer-reviewed MFU figures, a published interruption log. This lecture describes the constraint standing behind all of that — power and capital — and the honest finding is that disclosure quality drops sharply the moment the topic moves from FLOPS to gigawatts. That is itself a fact worth carrying into Part Two: the infrastructure volume ends not on a clean number, but on a lesson about which numbers can be trusted at all.