Capability Audit · Mainstream AI · Nov 2022 → Feb 2028

The Tier Ledger
Eight semesters scored, three projected.

Every benchmark saturates and dies; the frontier does not. This instrument normalizes eight domains onto a single human-anchored tier scale (0–5) and records three readings per domain, per half-year: the best system evidenced anywhere, including restricted access (frontier), the best model an individual can actually buy and wield (GA — your row), and what the broad economy reliably runs (deployed). Hatched cells are projections. Tier 5 is the top of the scale and is already open-ended — “beyond any human” — so a projection band that flattens against 5 means the scale has run out of room, not that the uncertainty has gone away.

What this is. A capability audit I keep for my own decisions: which model to build on, and what to expect the economy to be running two semesters from now. Compiled 16 August 2026 from the cited evidence; the 2026 H2 column turns from projection to measurement at the January 2027 refresh. The benchmarks in the evidence log are measurements; the tier number is not. Placing a result on a human-anchored scale is my judgement, made against that log and recorded so it can be argued with. Click a domain for its evidence log; the link in the address bar follows you. How the ledger changed →

01 · Composite

The frontier line and its shadow

Mean tier across all eight domains. The gap between the lines is the adoption lag — deployed capability in any given semester roughly equals benchmark capability from 2–3 semesters earlier. The red rule is today.

02 · The Ledger

Eight domains, eleven semesters, three truths each

F = frontier (best evidenced anywhere, incl. restricted access) · G = GA, the best you can buy — your row · D = deployed across the economy. Click a domain for its chart and evidence log. Hover any cell for the tier value. Columns right of the red rule are projected (center of band shown).

03 · Domain Detail

Mathematics

━ frontier   ━ GA (your row)   ━ deployed   ▒ projection band
04 · The Spread

Where the lab and the world disagree

Mid-2026, two gaps per domain. The wide amber bar is frontier − deployed — a liability measurement: license walls hold it open. The thin green bar is frontier − GA — the access gap, what restricted tiers withhold from an individual buyer. Note the scale difference: the economy runs 1–3 tiers behind the lab; you run about a tenth of a tier behind it.

05 · Findings

Rate

+0.43 tier / half-year

Composite frontier climbed from 0.9 to 3.9 in eight semesters — and the pace accelerated: METR's measured doubling time for autonomous task length compressed from ~7 months (2019–2025) to ~3.5–4 months (2025–26).

Lag

2–3 semesters

Deployed reality trails the lab by a year to eighteen months in every domain. Rule of thumb the ledger supports: the deployed world of 2027 is the benchmark world of 2025.

Wall

2.6 vs 0.4

Medicine's spread (2.6 tiers) vs writing's (0.4). Diagnostic AI beats physician panels on published cases while zero generative-AI devices hold FDA authorization. The spread prices regulation, not capability.

Projection

3 domains ≥ tier 5

By end-2027 (center case): math already past 5 — AI-primary solutions to open problems began 2026 — with code and science approaching. Agentic is the steepest slope; robotics the slowest. Bands widen with distance.

Access Gap

≈ 0.13 tier

Frontier minus GA, composite, today — roughly one release cycle. It breathes with the calendar: it widened to ~0.19 in 2025H2 (IMO-gold systems unreleased — math alone gapped 0.5 — restricted tiers forming) and narrowed as releases landed. The tell for it opening permanently: release cadence thinning while capex accelerates.

06 · Method & Caveats

How the ledger is built

Why tiers, not scores. No benchmark survives contact with the frontier: GSM8K → MATH → AIME → FrontierMath each saturated and was retired; SWE-bench Verified was deprecated by OpenAI in Feb 2026 over contamination after leaders passed 90%; OSWorld crossed its own human baseline in Dec 2025 and was immediately succeeded by a long-horizon version where the best system scores 20.6%. The tier scale chains successive benchmarks to a stable human anchor so semesters remain comparable.

Three rows per domain. The frontier row records the best credible published result anywhere, including restricted-access and vetted-tier systems (Mythos-class, gov-preview builds). The GA row records the best generally-available model an individual can buy at retail — subscription or API, no vetting — the row that defines a solo operator's actual toolkit. The deployed row records what works reliably in broad production — the number that moves payrolls. Historically GA ≈ frontier (labs shipped their best); the rows were split in 2025–26 by the access pyramid: internal → government preview → vetted partners → public. METR's cross-lab pilot (Feb–May 2026) found public capabilities still representative of the industry's strongest — the gap is real but currently months, not tiers.

Projection method. Per-domain tier-crossing cadence fitted over the eight observed semesters, sanity-checked against compute growth (~4–5×/yr), algorithmic efficiency (~3×/yr), and the METR horizon trend; three semesters projected with widening bands. Center case assumes no regime break in either direction — no takeoff, no wall, no post-incident freeze.

Caveats, honestly. Tier placement is judgment anchored to cited evidence — a different analyst shifts cells by ±0.3. Vendor self-reports inflate; contamination inflates; the METR logistic fit is fragile above ~12 hours (measurements above 16h flagged unreliable by METR itself). Deployed tiers lean on adoption surveys and regulatory records, which lag. Treat cells as calibrated estimates, not measurements.

Sources

07 · Version log

How the ledger changed

The ledger is refreshed every six months. Structural changes are recorded here; cell-level revisions are recorded in each domain's evidence log.

v1.016 Aug 2026

Eight domains, two rows each: the best benchmark result anywhere versus what the broad economy reliably runs. Eight semesters scored since ChatGPT, three projected with widening bands. First readings: composite benchmark 0.87 → 3.89 (+0.43 tier per half-year); deployed reality 2–3 semesters behind; medicine's spread 2.6 tiers against writing's 0.4.

What was wrong. It answered the wrong question for the person reading it. Economy-average deployment is the right number for a labour model; it is the wrong number for an operator who always runs the latest generally-available model. The row that matters to that reader — what can I buy today — did not exist.

v1.116 Aug 2026 · three lenses

Added the GA row between frontier and deployed: the best model an individual can buy at retail, no vetting. Splitting it from the frontier row made the access pyramid measurable: the frontier–GA gap is about 0.13 tier today, roughly one release cycle, and it peaked at 0.19 in 2025 H2 when IMO-gold systems went unreleased (mathematics alone gapped 0.5). The tell to watch for it opening permanently: release cadence thinning while capital expenditure accelerates.

v1.1 · webSep 2026 · publication

Published unchanged in substance. Added: the selected domain encoded in the URL, a copy-link button, and this log. The next refresh (January 2027) will grade the 2026 H2 projections against what was actually measured, and record the misses.

Link copied