Faster than Moore.
For fifty years, Moore's law was the fastest sustained exponential in the industrial record: transistors per chip doubled every — months, from 2,308 in 1971 to 58 billion. AI training compute tracked that same slope for six decades — and then, around 2010, it left. Measured across — AI models with a documented training run, compute has doubled every — months since 2010 — roughly four times faster than the law that built the chips it runs on — reaching 5×10²⁶ floating-point operations for the largest run on record. Every slope on this page is fitted from the records at render time.
published 2026-08-27 · data cutoff 2026-08 (Epoch AI database, latest) · edition 1 — updated only by a new dated edition · source: Epoch AI · OWID/Karl Rupp
Seventy-six years, twenty-four orders of magnitude.
Every AI model with a documented training run since 1950, on one log field — each gridline is a thousandfold. The gray line is Moore's law itself, drawn from the transistor record. For sixty years the dots climb with the line. Around 2010 they break upward and never come back — the fitted slopes are drawn and labeled, and the gap between them is this report's finding.
The 10²⁵ club.
The largest documented training runs, ranked. Grok 4 sits at 5×10²⁶ floating-point operations — roughly a billion times the compute of AlexNet, the 2012 model that opened this era. Everything at this scale is younger than 2023, and the club above 10²⁵ has — members.
Who holds the frontier.
The 10²⁵ club by organization and by country. The frontier of compute is held by a handful of hands — and nearly all of the club's members were trained in the United States, with China's largest runs an order of magnitude behind. Concentration is the frontier's second signature, after speed.
Moore's law never bent.
The transistor record itself, 1971–2021. The remarkable thing is its steadiness: a —-month doubling held for fifty years. The escape of 2010 did not come from the chips suddenly improving faster — it came from money and parallelism: more chips, bigger clusters, longer runs. The capital side of that story is The Data Center Decade; the electricity it draws is The Power Mix.
Where every number comes from.
Two records. Epoch AI's Notable AI Models database (CC-BY) — every model with a documented or estimated training compute in FLOP, with organization, country, domain and Epoch's own confidence grade carried in the record. And the transistor-count series compiled by Karl Rupp and published by Our World in Data, 1971–2021. Both fetched 2026-08-27; both ship in data.json beside this page.
The slopes are fitted in your browser — least-squares on log₁₀ compute against time, one fit for the era before 2010 and one after, and the same fit on the transistor series. A doubling time is 12·log₁₀(2) divided by the slope, in months. Change the record and the slopes change with it.
What a FLOP count is. Training compute is documented for some models and estimated by Epoch for others; the record carries a confidence grade per row and this page inherits it rather than re-judging it. The 2010 break year follows the standard periodization in the literature; the fit does not depend on the exact year chosen.
The overlay is a slope comparison. Transistors-per-chip and training FLOP are different quantities; the gray line is scaled to the field so their rates can be compared. No unit equivalence is claimed anywhere on this page.
What this page does not do. No capability claims, no forecasts, no scaling-law extrapolation. The record's own arithmetic — slopes, ranks, counts — is the entire method.