🔭 Futures

AI agent autonomously handles 30%+ of a typical knowledge-work job

conf: medium
P50 pushed out by 1 year
2026-06-16 → 2026-09-07
P502028 → 2029 (P10 2027 → 2028, P90 2033 → 2033)

Reverts the June pull-in: both legs of it failed re-verification, and the first direct measurement of the gate's own criterion landed below the bar. (1) METR — the dashboard's SOTA has been frozen at 17.41h (Claude Mythos Preview, released 2026-04-07, added 2026-05-08) with no new model added in four months, above METR's own posted ceiling 'Measurements above 16 hrs are unreliable with our current task suite'; the published doubling estimate is now 128.7d (2023-on, CI 104-158) and 187.8d all-time-stitched, not the 89d this gate assumed, and METR pivoted to a new 'expenditure horizon' metric in July 2026. The gate's own METR dependency prices a doubling-rate reversion at +1-2y; taking the low end gives +1y. (2) SWE-bench Pro — Scale's standardized public set (identical scaffolding) tops out at 61.5% (Muse Spark 1.1), not the 77.8/80.3% vendor-harness figures the June entry read off it, so the contamination-resistant 75% sub-gate is not cleared. (3) Anthropic's 2026 Agentic Coding Trends Report measures the criterion directly: engineers use AI in ~60% of their work but can 'fully delegate' only 0-20% of tasks; Anthropic's own 'When AI builds itself' reports >80% of merged code Claude-authored while naming human code review the new bottleneck — the 'without human review on each step' clause is exactly what has not happened. GPT-6 Astra (2026-09-03) is a real capability jump but ranks 22nd on GDPval-AA v2 (1,582 vs Fable 5.1's 1,766) and 6th on AA-Briefcase, and OpenAI omitted GDPval from launch materials. Confidence stays medium: capability and commercial momentum are intact, so P90 holds at 2033. — https://metr.org/time-horizons/; https://labs.scale.com/leaderboard/swe_bench_pro_public; https://resources.anthropic.com/2026-agentic-coding-trends-report; https://www.anthropic.com/institute/recursive-self-improvement; https://artificialanalysis.ai/evaluations/gdpval-aa

Trigger
An AI agent autonomously completes ≥ 30% of tasks in a typical knowledge-work role (e.g. legal review / software dev / content / customer support) without human review on each step. 'Typical' = the modal task distribution of the role per O*NET or comparable taxonomy.
Timeline
2027
2030
2033
2036
2040
2045
2050
P10 2028
P50 2029
P90 2033
50 sources last updated: 2026-09-07 View raw .md ↗
Prediction history
3 entries · latest first
  1. 2026-09-07
    current
    P10 2028 · P50 2029 · P90 2033
    Reverts the June pull-in: both legs of it failed re-verification, and the first direct measurement of the gate's own criterion landed below the bar. (1) METR — the dashboard's SOTA has been frozen at 17.41h (Claude Mythos Preview, released 2026-04-07, added 2026-05-08) with no new model added in four months, above METR's own posted ceiling 'Measurements above 16 hrs are unreliable with our current task suite'; the published doubling estimate is now 128.7d (2023-on, CI 104-158) and 187.8d all-time-stitched, not the 89d this gate assumed, and METR pivoted to a new 'expenditure horizon' metric in July 2026. The gate's own METR dependency prices a doubling-rate reversion at +1-2y; taking the low end gives +1y. (2) SWE-bench Pro — Scale's standardized public set (identical scaffolding) tops out at 61.5% (Muse Spark 1.1), not the 77.8/80.3% vendor-harness figures the June entry read off it, so the contamination-resistant 75% sub-gate is not cleared. (3) Anthropic's 2026 Agentic Coding Trends Report measures the criterion directly: engineers use AI in ~60% of their work but can 'fully delegate' only 0-20% of tasks; Anthropic's own 'When AI builds itself' reports >80% of merged code Claude-authored while naming human code review the new bottleneck — the 'without human review on each step' clause is exactly what has not happened. GPT-6 Astra (2026-09-03) is a real capability jump but ranks 22nd on GDPval-AA v2 (1,582 vs Fable 5.1's 1,766) and 6th on AA-Briefcase, and OpenAI omitted GDPval from launch materials. Confidence stays medium: capability and commercial momentum are intact, so P90 holds at 2033. — https://metr.org/time-horizons/; https://labs.scale.com/leaderboard/swe_bench_pro_public; https://resources.anthropic.com/2026-agentic-coding-trends-report; https://www.anthropic.com/institute/recursive-self-improvement; https://artificialanalysis.ai/evaluations/gdpval-aa
  2. 2026-06-16
    P10 2027 · P50 2028 · P90 2033
    SWE-bench Pro (Scale AI, contamination-resistant) cleared: Claude Mythos Preview 77.8% / Claude Fable 5 80.3% as of June 2026 — sub-gate swe-bench-pro-75pct resolves ~1yr ahead of schedule. METR 50% time horizon confirmed ~14h (public models, Feb 2026) with ~7-month doubling rate sustaining → 40h crossing (metr-time-horizon-1-week) on track for early 2027, ~1yr ahead of prior P50 2028. Both binding sub-gates shifting 1yr earlier cascades main P50 from 2029→2028. — https://labs.scale.com/leaderboard/swe_bench_pro_public; https://metr.org/blog/2026-1-29-time-horizon-1-1/
  3. 2026-05-13
    P10 2027 · P50 2029 · P90 2034
    Initial estimate from initial research.
Key dependencies — watch these
  • METR doubling rate sustains or reverts both
    PARTIALLY FLIPPED Sept 2026 and worth 1 of the priced 1-2 years. METR's published doubling estimate is now 128.7 days (2023-on, CI 104-158) and 187.8 days all-time-stitched, not the 89 days this gate assumed; the dashboard SOTA has been static at 17.41h (Claude Mythos Preview) since 2026-05-08 with no new model added. A full reversion to the ~7-month all-time rate from here would cost another ~1 year; a resumed sub-100-day rate confirmed on a re-scaled task suite would return P50 to 2028.
  • METR time-horizon instrument saturation delays
    METR posts that 'measurements above 16 hrs are unreliable with our current task suite' and the SOTA point estimate (17.41h) already sits above it with a 95% CI of 8.5-55h; METR shipped a different metric ('expenditure horizon') in July 2026 rather than extending the suite. Until a longer-task suite ships, the 40h sub-gate cannot resolve either way, which caps confidence at medium regardless of underlying capability.
  • §
    EU AI Act high-risk agentic ban delays
    Deadline now settled in law, not proposed: the Digital Omnibus on AI (Regulation (EU) 2026/1744) entered into force 2026-07-27, deferring standalone Annex III high-risk obligations to 2027-12-02 (Annex I to 2028-08-02) while expanding AI Office inspection powers. A hard unsupervised-action reading at that deadline carves out 30-40% of EU knowledge-work roles and slows US adoption by precedent, dropping confidence one step.
  • Training data exhaustion 2027-28 delays
    If synthetic data and RL-from-environment fail to compensate for data limits, the doubling rate breaks and P50 slides 3+ years to ~2032. Not yet flipped — frontier capability kept climbing through 2026 (ARC-AGI-3 and FrontierMath Tier 4 both effectively saturated by GPT-6 Astra) — but METR's July 2026 expenditure-horizon result (best agents ~$2-3K crossover vs ~$2,500 per 1% of human labor on NanoGPT, 'minimal effect on AI R&D progress') is the first quantitative sign that the self-acceleration channel is not yet compensating.
  • Enterprise integration friction vs capability gap accelerates
    Resolving toward the slow branch: ~88% of enterprise agent pilots still fail to graduate to production (<15% of pilot-running enterprises reach it), blockers cited as evaluation gaps 64% / governance 57% / model reliability 51%, and the Anthropic Economic Index finds measurable Claude usage on only 7.5% of 17,998 O*NET tasks. The 'plumbing alone gets there' branch that would have opened P10 to 2027 is not the branch we are on; a reversal — pilot-to-production conversion above ~40% — would pull P50 back to 2028.
  • §
    California Bar and sectoral AI oversight rules delays
    Hardened from ethics proposal to statute: California SB 574 passed the legislature 2026-08-31, requiring attorneys to personally read and verify every AI-cited source, disclose AI use in court filings, and barring delegation of the practice of law to AI — extended to arbitrators. Together with the State Bar's May 2026 Rule 1.1 / 5.1 / 5.3 amendments covering agentic tools, mandatory human-verification in legal, finance and HR blocks the autonomous-without-review threshold in regulated roles, delaying sectoral P50 by 2+ years.
  • Unmonitored task-share measurement exists at all delays
    The gate resolves on a measured O*NET-referenced share of tasks done without per-step review, and no such measurement program exists. The closest direct read is Anthropic's 2026 Agentic Coding Trends Report — engineers use AI in ~60% of their work but can 'fully delegate' only 0-20% of tasks — i.e. below the 30% bar in the most AI-saturated role. A credible published role-level audit crossing 30% would move P50 in by 1-2 years on its own; continued absence of any instrument holds P50 at 2029 or later.
  • Shared long-horizon tool-using agent stack means crossing 30% in knowledge work essentially confirms K-8 parity is a packaging problem; the Sept 2026 slip of this gate's P50 to 2029 pushes the K-8 knock-on out by the same ~1 year, and a re-acceleration to 2028 would pull K-8 P50 in by ~6 months.
  • 30%+ knowledge-work substitution is the most direct upstream driver of sustained 10% US unemployment; faster-than-P50 crossing compresses the unemployment gate's P10 by 2-3 years. Currently pointing the other way — Challenger counted 529,914 announced US job cuts Jan-Aug 2026, down 41% year-on-year and the lowest January-August total since 2022, with AI falling to the fourth-most-cited reason in August (3,462).

Why this refresh moved the timeline (September 2026)

GPT-6 Astra (2026-09-03) triggered this refresh, but it is not what moved it. Astra is a genuine capability jump — ARC-AGI-3 at 99.9%, FrontierMath Tier 4 at 97.6%, OSWorld 2.0 at 72.6% in ~40 min/task against Sol’s 65.7% at ~75 min, a 1.05M-token window with 100% MRCR v2 at 256K-512K — and Greg Brockman used the launch to declare “we are now in the AGI era” [30][31][32]. It is also mid-pack on exactly the work this gate is about: 22nd on Artificial Analysis’s GDPval-AA v2 (1,582 Elo at max effort against Claude Fable 5.1’s 1,766) and 6th on AA-Briefcase’s long-horizon knowledge-work suite, and OpenAI omitted GDPval — its own benchmark for economically valuable real-world tasks — from the launch materials entirely [33][34][35]. That is a frontier-model release, which this gate already prices in.

What moved the timeline is that both legs of the June 2026 pull-in failed re-verification, and the first direct measurement of this gate’s actual criterion came in below the bar. (1) METR’s dashboard SOTA has been frozen at 17.41 hours since Claude Mythos Preview was added on 2026-05-08 — no new model in four months — and that point estimate already sits above METR’s own posted warning that “measurements above 16 hrs are unreliable with our current task suite,” with a 95% CI of 8.5-55h. The doubling estimate METR now publishes is 128.7 days (2023-on, CI 104-158) and 187.8 days all-time-stitched, not the 89 days this gate assumed; in July 2026 METR shipped a different metric (“expenditure horizon”) rather than extending the task suite [36][37]. (2) Scale’s standardized SWE-bench Pro public leaderboard — every model through identical scaffolding — tops out at 61.5% (Muse Spark 1.1), with GPT-5.4 xHigh at 59.1%; the 77.8%/80.3% figures the June entry read as clearing the 75% sub-gate are vendor-reported on different harnesses [38]. (3) Anthropic’s 2026 Agentic Coding Trends Report measures the criterion head-on: engineers use AI in ~60% of their work but can “fully delegate” only 0-20% of tasks [39]. Anthropic’s own When AI builds itself reports >80% of merged production code authored by Claude — and names human code review as the new bottleneck, with automated review catching about a third of bugs that reached production [40]. The “without human review on each step” clause is precisely what has not happened at the most AI-saturated employer in the world.

The gate’s own METR dependency prices a doubling-rate reversion at +1-2 years. Taking the low end — because this is partly instrument saturation rather than demonstrated capability stall — gives +1 year: P50 2028 → 2029, P10 2027 → 2028. P90 stays 2033 and confidence stays medium: capability and commercial momentum are entirely intact (Cognition went from $492M to >$900M ARR between May and September 2026 at a reported $47B valuation, Cursor is at $2B, Anthropic’s internal open-ended-task success went 26% → 76% in six months), and none of the new evidence lengthens the tail.

TL;DR

I put the P50 at 2029 — pushed back from 2028 in September 2026 when the two sub-gate claims behind the June pull-in did not survive re-checking — that an AI agent will autonomously complete ≥30% of tasks in at least one typical knowledge-work role, measured against an ONET-style task inventory and without per-step human review. The headline thesis has not changed: customer support is the closest to crossing (Klarna handled 67% of conversations end-to-end, though it re-hired humans for the hard tail and now describes the hybrid as the destination rather than a retreat), and software engineering is visibly mid-cross with >80% of Anthropic’s merged code Claude-authored and Devin writing 89% of Cognition’s own commits. What the 2026 evidence sharpened is which gap is binding. It is not raw capability. It is the difference between assisted throughput (large, growing, measurable) and unreviewed task share (0-20% by the only direct survey we have), plus integration friction — ~88% of enterprise agent pilots never reach production — and policy and liability, which hardened rather than softened in 2026. P10 = 2028 (one role, a fast-moving enterprise, formally audited against ONET); P90 = 2033 (a real stall in agent reliability or a hard regulatory clamp pushes the threshold out a decade). The single most important quantitative driver is still the METR time horizon — but as of September 2026 that instrument has stopped resolving, which is itself the news.

Current state (as of 2026-09-07)

The 30% threshold is still not crossed in any agreed-upon “typical knowledge-work job” measured against O*NET, and the September 2026 evidence tightened rather than loosened that judgment. Four hard numbers anchor where we are:

  • METR time horizon: SOTA is 17.41 hours (Claude Mythos Preview, released 2026-04-07, added to the dashboard 2026-05-08) with a 95% CI of 8.5-55h — above METR’s own stated reliability ceiling of 16h, and unchanged for four months despite GPT-5.6 Sol, Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra all shipping since. Published doubling: 128.7 days from 2023 on (CI 104-158), 187.8 days all-time-stitched. A June 2026 pre-deployment eval of GPT-5.6 Sol returned only ~11.3h under standard scoring [36][37].
  • Delegation share: Anthropic’s 2026 Agentic Coding Trends Report — engineers use AI in ~60% of their work but can fully delegate only 0-20% of tasks; ~27% of AI-assisted work is work that would not have been attempted at all [39]. Anthropic internally: >80% of merged production code Claude-authored as of May 2026 (from low single digits pre-Claude-Code), engineers merging 8× the code per day versus 2024, open-ended internal task success 26% (Nov 2025) → 76% (May 2026) — with the caveat, from Anthropic, that lines-of-code “almost certainly overstates the true productivity gain” and that human review is now the bottleneck [40].
  • Benchmarks: Scale’s standardized SWE-bench Pro tops at 61.5% (Muse Spark 1.1) [38]. GDPval-AA v2 (44 occupations, 9 sectors, agentic environments, LLM-judged pairwise Elo against a 1,000 human baseline): Claude Fable 5.1 1,766, Claude Opus 5 1,738, GPT-6 Astra 1,582 [33]. AutomationBench (Zapier; 657 cross-application business-workflow tasks across finance, HR, marketing, operations, sales, support): best score 41.4% [32]. τ²-bench clears 90% on Artificial Analysis’s harness (99.1%) but ~88% on vendor-mixed aggregates. OSWorld 2.0: 72.6% [30][31].
  • Deployment and labor market: Cognition at >$900M ARR (from $492M in May 2026) with 89% of its own code written by Devin and a reported $47B round; Cursor at $2B ARR [41]. Against that: ~88% of enterprise agent pilots fail to reach production, with evaluation gaps (64%), governance friction (57%) and model reliability (51%) as blockers [42]; the Anthropic Economic Index finds measurable Claude usage on only 7.5% of 17,998 O*NET tasks [43]; and Challenger counted 529,914 announced US job cuts January-August 2026, down 41% year-on-year and the lowest Jan-Aug total since 2022, with AI falling to the fourth-most-cited reason in August (3,462 cuts) [44].

So as of September 2026: capability keeps climbing, deployment revenue keeps compounding, and the unreviewed share of a real role remains the thing nobody can yet show above 30%. The lagging variable is still measurement — and it got worse, not better, this quarter: METR’s series stopped resolving, Scale’s standardized SWE-bench Pro set stopped receiving frontier models, and OpenAI declared the AGI era while withholding its own economic-work benchmark.

Key uncertainties

  1. Has the METR doubling rate actually slowed, or has the instrument simply saturated? METR’s published estimate moved from 89 days (2024-on, TH 1.1 blog) to 128.7 days (2023-on, current dashboard), and the SOTA point has been static at 17.41h for four months — but METR also says anything above 16h is unreliable on the current suite and has not added the four frontier models released since. Extrapolating 17.41h at 128.7-day doubling puts 40h at roughly now; we simply cannot confirm it. Whether a re-scaled suite shows the trend intact or broken is the single highest-value observation for this gate.
  2. What % of tasks in a “typical” job are within agent capability today but blocked by integration / context / data access? The 2026 read is that plumbing and governance dominate: ~88% of pilots die before production and only 7.5% of O*NET tasks show measurable frontier-model usage at all. That is the slow branch. Pilot-to-production conversion above ~40% would flip it.
  3. Can silent failure be monitored without a human in the loop? This is the load-bearing question for “without per-step review” and it got a real answer in 2026: 45-48% of τ²-bench failures and 75.8% of AppWorld self-assessing coding-agent failures are false successes, and no LLM-judge configuration exceeds AUROC 0.65 at detecting them [45]. Cheap TF-IDF detectors do far better (0.83 / 0.95), which suggests the problem is tractable — but until unsupervised deployment has a monitor it can trust, per-step review is the rational default regardless of capability.
  4. Does the EU AI Act’s December 2027 high-risk deadline functionally ban unsupervised agentic action in regulated domains (legal, finance, HR)? No longer speculative on timing: the Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force 2026-07-27 — deferring Annex III high-risk obligations to 2027-12-02 and Annex I to 2028-08-02, while expanding the AI Office’s inspection and binding-commitment powers [46]. Deferred, not cancelled, and the enforcement side got stronger.
  5. Does training data run out in 2027-28 and stall scaling? Not yet visible in capability — ARC-AGI-3 and FrontierMath Tier 4 were effectively saturated in 2026 — but METR’s expenditure-horizon result is the first quantitative hint that the self-acceleration channel isn’t compensating: on the NanoGPT speedrun the best agents’ crossover sits at ~$2-3K against ~$2,500 of human labor per 1% improvement, and METR concludes autonomous agent optimization “has so far had minimal effect on AI R&D progress” there [37].
  6. Will labor pushback and sectoral rules make 30%-autonomous a no-go even if the tech works? Moving from rhetorical to statutory: California SB 574 passed the legislature 2026-08-31, requiring attorneys to personally verify every AI-cited source, disclose AI use in filings, and not delegate the practice of law to AI — extended to arbitrators [47].

Evidence synthesis

Academic

The single most important academic anchor is METR’s Measuring AI Ability to Complete Long Software Tasks (Kwa et al., arXiv:2503.14499) [11], which introduces the time-horizon metric: the duration of human-expert task that the model can complete with 50% reliability. METR’s longitudinal data shows a 7-month doubling from 2019–2025 (~9 seconds for GPT-3 in 2020 → ~50 minutes for Claude 3.7 in early 2025). The January 2026 TH 1.1 blog reported doubling compressing to ~89 days since 2024 [2], which is the figure this gate ran on through June 2026.

That figure no longer matches what METR publishes. As of September 2026 the live dashboard’s own embedded data gives doubling of 128.7 days from 2023 on (95% CI 104.4–158.0) and 187.8 days all-time-stitched, and the model series under TH 1.1 reads: Claude Opus 4.5 293 min (Nov 2025) → GPT-5.2 352 min (Dec 2025) → Claude Opus 4.6 719 min ≈ 12.0h (Feb 2026, CI 5.3–60.6h) → Claude Mythos Preview (early) 1,045 min ≈ 17.41h (released 2026-04-07, added 2026-05-08, CI 8.5–55.1h). Mythos is still flagged is_sota four months later: GPT-5.6 Sol, Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra have not been added. METR’s own chart annotation reads “Measurements above 16 hrs are unreliable with our current task suite” — i.e. the SOTA point estimate sits above the stated reliability ceiling, and a June 2026 pre-deployment evaluation of GPT-5.6 Sol returned only ~11.3h under standard scoring, with METR noting heavy sensitivity to reward-hacking behaviour [36]. In July 2026 METR published a different metric rather than a longer suite: the expenditure horizon, the dollar budget at which humans become more cost-effective than agents on continuously-scored optimisation. On the NanoGPT speedrun, the best models (GPT-5.5, Opus-4.8) reach crossover at roughly $2.3K and $3.3K against an estimated $2,500 of human labour per 1% improvement — i.e. about parity — and METR’s own summary is that “autonomous agent optimization has so far had minimal effect on AI R&D progress” there [37]. Naive extrapolation from 17.41h at 128.7-day doubling puts the 40h threshold at roughly now; the defensible statement is that the instrument stopped resolving before it got there.

On benchmark-specific results, the SWE-bench family (Princeton/OpenAI, agent code repair) and its harder variants (SWE-bench Pro [4], SWE-bench Multimodal) show the same exponential — but with explicit warnings that Verified saturation is partly contamination [3]. The GAIA benchmark (Meta AI Research) for general AI assistants shows frontier models at 74.6% (Claude Sonnet 4.5, scaffolded) vs human baseline of ~92% [12]. OSWorld for computer-use agents has now been crossed by GPT-5.5 (78.7% vs ~72% human) [5]. τ²-Bench (Sierra) measures customer-service agents on policy-adherent multi-turn flows — a much harder bar than “task complete” — and frontier models cluster in the 50–65% range as of April 2026 [13].

Citation-graph leaders in agent eval are Princeton (SWE-bench, HAL), Sierra (τ-bench), CMU+OSU (WebArena, OSWorld), and METR itself. The trajectory across all five benchmarks is consistent: rapid progress on narrow tasks (coding), slower on broad open-world tasks (GAIA, OSWorld), and the slowest on multi-turn-with-policy (τ-bench) — which is also the most predictive of real deployment success.

Anthropic’s own Project Vend [14] is informative as a negative result: in mid-2025 Claude failed to profitably run a tiny vending business autonomously over a month; in Phase 2 (with multi-agent architecture) it improved meaningfully but the gap between “capable” and “completely robust” remained wide. This is good evidence that agentic business operation is harder than agentic task completion, and is why I don’t think benchmark scores alone get us to a 30%-of-a-job claim.

Mid-2026 update — the reliability literature caught up with the capability literature. Three post-June papers bear directly on whether per-step human review can be removed, and all three point the same way. From Confident Closing to Silent Failure (Advani, arXiv:2606.09863, FAGEN@ICML2026) characterises false success — the agent asserting completion when the environment state says otherwise — across 9,876 τ²-bench trajectories from 8 model families and 1,879 AppWorld trajectories from 4: 45-48% of failures in single-control τ²-bench domains, 3% in dual-control telecom, 75.8% among AppWorld self-assessing coding agents. No LLM-judge configuration (5 judges × 5 prompt strategies, full task specs) exceeds AUROC 0.65; judges key on “confident closing language” rather than verified state changes. Lightweight TF-IDF detectors reach 0.83/0.95 at 3,300× lower latency, so this is fixable — but today the default monitor for an unsupervised agent cannot see half its failures [45]. The Horizon Gap (Chen et al., arXiv:2608.06663, 7 Aug 2026) surveys 1,547 papers and names the pattern: outcome-only signals grow uninformative as horizons lengthen, with coherence collapse and goal drift — including compression mechanisms that “silently erase the safety constraints” given at trajectory start — and notes SWE-bench resolution for one leaderboard-topping system falling from 12.47% to 3.97% after contamination and weak-test filtering [48]. Deployment Decision Reliability (Srinivasan, arXiv:2608.11323, 11 Aug 2026) decomposes variance across TheAgentCompany, τ²-bench and AppWorld and finds the agent main effect explains <3% of total variance while agent-by-task interaction explains 7-23%: “leaderboards rank specialization, not capability.” Aggregate reliability collapses on the hardest task quartile (Eρ² on τ² action_checks falls 0.752 → 0.000), and training-cell reliability negatively correlates with held-out reliability (r = −0.90) — the designs that look most reliable replicate worst [49].

That literature is the reason I read GDPval-AA Elo and τ²-bench scores as necessary but nowhere near sufficient. It is also consistent with the Princeton-led open-ended-research evaluation reported in August 2026, where Claude Opus 4.8 agents given six days and $3,000 to attack unpublished NeurIPS 2026 questions were “capable of all the engineering required” but “unambiguously bad at carrying out the research itself” — both agent-generated papers were rejected by the original authors grading them [50].

On benchmark supply: TheAgentCompany (CMU) remains the closest public analogue to this gate’s framing — a simulated software company where the best agents autonomously complete ~30% of tasks (39.3% with partial credit) — and its authors’ conclusion still stands: agents “are not close to automating every task encountered in a workspace, even on the subset presented,” with social interaction, professional UIs, and tasks lacking public documentation the hardest. Two newer suites are more directly on-point and both leave large headroom: AutomationBench (Zapier; 657 tasks across 47 simulated SaaS tools in finance, HR, marketing, operations, sales and support, scored on final environment state with no LLM judge) has a best score of 41.4%, and AA-Briefcase (91 tasks across four professional scenarios spanning up to six simulated weeks, producing spreadsheets/decks/memos) is led by Claude Fable 5.1 with GPT-6 Astra sixth [32][34].

Industry / market

The deployment numbers from 2025–2026 are the strongest single piece of evidence that we are mid-crossing, not approaching, this gate. Five anchor points:

  1. Anthropic crossed $30B annualized revenue by April 2026, up from $9B at end-2025 and $1B at end-2024 — an 80x growth in 16 months [9]. Claude Code alone is at $2.5B run-rate, with enterprise representing >50% of Claude Code revenue. >1,000 customers now spend >$1M annually with Anthropic, vs ~12 two years earlier.
  2. Cursor hit $2B ARR by February 2026 from $100M one year earlier — used by ~70% of the Fortune 1000 [6]. Latest reporting puts Cursor in talks at a $50B valuation.
  3. Cognition / Devin: PR merge rate up to 67% in the 2025 performance review (from 34% the prior year); deployment at thousands of companies including Goldman Sachs, Santander, Dell, Cisco, Palantir; Infosys partnership for global financial-services rollout being “one of the largest agentic deployments to date” [7]. ARR went from $1M (Sep 2024) to $73M (Jun 2025). January 2026 launched Devin Review for automated code review.
  4. Replit at $150M ARR (Sep 2025), on track for $1B run-rate by end-2026 [15]. Lovable at $200M ARR (Nov 2025), 8M users, 100K+ projects/day [15]. Both target the “non-engineer builds an app” segment which is itself a knowledge-work automation play.
  5. Klarna: in customer support specifically, the 67% autonomous resolution rate (Feb 2024 launch month) achieved a 40% drop in cost-per-transaction over 2 years and replaced ~700 agents [10]. Note the 2025 walkback: Klarna re-introduced humans for the 5% of complex/emotional/edge-case conversations. This is the clearest existing example of crossing the 30% threshold in a real role today — though customer support is the easiest knowledge-work role to automate.

The Anthropic Economic Index (March 2026) [16] is the best public dataset on actual usage patterns: API traffic (which is more agentic / automated) shows the share of “computer and mathematical” tasks growing +14% over 6 months, and customer service is flagged as the occupational category with highest automation exposure. New API automation categories emerging in late 2025 include “business sales and outreach automation” and “automated trading and market operations” — both >2x growth. Critically, Anthropic notes that automation already dominates 1P API traffic, meaning that for the workloads developers ship to customers, the median interaction is already minimal-human-loop.

McKinsey Global Institute’s 2023 estimate — reaffirmed in 2025 — pegs up to 30% of US work hours automatable by 2030 with genAI [17]. For STEM specifically, McKinsey raises the 2030 automation potential from 14% to 30% under genAI scenarios. Forrester’s 2024 estimate is more conservative (6% of US jobs eliminated by 2030), reflecting the distinction between task automation and job elimination that explains why my gate is well-defined: 30% of tasks within a role is a much earlier milestone than 30% of jobs gone.

Mid-2026 update — the market got much bigger while the delegated share stayed small. Cognition went from $492M run-rate revenue in May 2026 to >$900M by September, with 89% of its own committed code written by Devin and a reported round valuing it at $47B (up from $26B in May); Cursor’s $2B ARR now carries a reported $60B acquisition option [41]. Anthropic’s When AI builds itself (June 2026) is the single richest internal dataset anyone has published: >80% of code merged into Anthropic’s production codebase in May 2026 was authored by Claude (from low single digits before Claude Code launched in Feb 2025); the typical engineer merges 8× as much code per day as in 2024; success on the hardest, least-specified internal coding tasks went 26% (Nov 2025) → 76% (May 2026); and on 129 selected research-detour moments Mythos Preview beat the human’s next-step choice 64% of the time, up from 51% for Opus 4.5 [40]. Anthropic supplies its own deflators, which matter for this gate: lines-of-code “measures quantity over quality… almost certainly an overstatement,” the 4× self-reported productivity median from a March 2026 survey of 130 staff is probably high (they cite METR’s own work on developers overestimating AI uplift), and — decisively — “human code review has become a new bottleneck,” with humans shifting from writing code to only reviewing it. Automated Claude review would have caught roughly a third of the bugs behind past production incidents; the other two thirds are why review persists.

Set against that, the deployment funnel is the constraint. Roughly 88% of enterprise agent pilots never graduate to production, and a March 2026 survey found 78% of enterprises running agent pilots with under 15% reaching production; the cited blockers are evaluation gaps (64%), governance friction (57%) and model reliability (51%) [42]. MIT’s NANDA review of 300+ disclosed deployments put 95% of enterprise genAI pilots at zero measurable return. The Anthropic Economic Index (June 2026, “Cadences”) is the cleanest ONET-referenced read: measurable Claude usage appears on only **7.5% of 17,998 ONET tasks**, knowledge work clusters at the top of adoption with physical work near zero, and the respondent base is heavily skewed (Computer & Mathematical occupations are ~30% of respondents against a 4% share of US employment) [43]. And the labor-market signal cooled: Challenger counted 529,914 announced US job cuts January-August 2026, down 41% year-on-year and the lowest Jan-Aug total since 2022, with AI slipping to the fourth-most-cited reason in August at 3,462 cuts, behind restructuring [44]. Klarna, still the clearest role-level example, now describes its hybrid — AI handling the two-thirds it handles well, humans rebuilt for the emotional and ambiguous tail — as the working end state rather than a retreat [10].

Public sentiment

r/cscareerquestions in May 2026 is dominated by layoff posts and AI-displacement narratives. The top post — “4 engineers now doing the job of 12 at my friend’s company because AI agents handle the rest” — has 1,315 upvotes and 516 comments [18]. The framing in earnings calls Microsoft, Meta, Cloudflare, Cisco, Coinbase, Paypal is consistent: “AI-driven efficiencies” justifying headcount cuts. CS enrollment is dropping for the first time in 6 years while ME and EE rise 11%/14% [18]. A senior FAANG manager’s hopeful counter-narrative (“the golden age is coming, you’re not cooked”) got 1,354 upvotes — but the top comments push back hard. Sentiment among working software engineers is directly observable as bearish on entry-level survival, bullish on senior-with-AI productivity. The “first rung of the ladder turns into a button, eventually there’s no ladder” meme has 1,602 upvotes.

r/LocalLLaMA shows a different signal: builders are skeptical of frontier-only narratives. Qwen 3.6 27B replacing Opus 4.7 for “85% of tasks” [19] suggests open-weight models are catching up fast — meaning the agent stack is commoditizing, not concentrating. This is bullish for “30% of role” timelines because it implies cost-per-task is falling fast enough to make economically viable agent deployment practical even for low-margin roles.

r/Lawyertalk is the most informative for legal-as-a-knowledge-work-role. Top post: clients showed up to an estate-planning consult wearing Meta camera glasses, had AI analyze the meeting, then sent back an AI-generated critique recommending an offshore trust and undercutting the lawyer’s fees [20]. The DOJ-brief-clearly-written-by-AI post had 1,266 upvotes. Lawyers are simultaneously contemptuous of AI legal output quality and aware it’s pricing into client expectations. California Bar’s May 2026 rule proposal requires verification of every AI output [21] — this is exactly the friction that delays “30% autonomous” in regulated knowledge work even when the capability is there.

r/singularity is, predictably, on the bullish extreme — the modal post in April–May 2026 is robot manufacturing acceleration, Atlas tricks, half-marathon records broken by robots, GPT-5.4 solving 60-year-old Erdős problems [22]. Sam Altman publicly walked back UBI advocacy (“I no longer believe in universal basic income as much as I once did”) — bullish signal that even OpenAI’s CEO no longer thinks the labor displacement will be smoothed by cash transfers, which is itself an admission of expected speed.

Prediction markets

The directly relevant Metaculus question is AI as a Competent Programmer Before 2030 [23] — community resolution implies a median probability close to 70–80% of competent autonomous programming by end-2029 under reasonable interpretations. The associated Metaculus “AI 2027” tournament aggregates questions on automation of AI R&D, with community estimates that AI begins automating AI research by 2027 — directly upstream of my gate. The Metaculus Labor Automation Forecasting Hub [24] hosts multiple questions on hours automated; the modal community estimate is consistent with McKinsey’s 30%-by-2030 number.

On Manifold, the cleanest question I found is “Will AI cause the US Unemployment Rate to exceed 10% before 2030?” [25] — priced at 25% as of September 2026 across 161 holders, up from the 15–20% recorded in May. Note the strict resolution bar: AI must be directly responsible for at least 3pp of the increase. This still sits well below any “AI handles 30% of tasks” question because task automation doesn’t translate 1:1 to unemployment — many displaced workers shift roles rather than become unemployed.

The Metaculus Labor Automation Forecasting Hub [24] is now the better calibration anchor than the single programmer question, and its numbers are strikingly structural rather than dramatic: forecasters expect US employment to fall ~3% by 2035 where BLS projects +3% growth (a six-point gap they attribute to AI); the most AI-exposed occupations shrink 17.2% while the least-exposed grow 4.6%; labour’s share of national income drops 62.1% → 55.6%; long-term unemployment rises 0.99% → 6.14%; young-graduate unemployment doubles 6% → 12%; and weekly hours fall 38 → 34 while wages still grow. Software developers and financial specialists are the largest projected decliners.

My P50 of 2029 now sits roughly on the Metaculus AI-programmer-by-2030 implied resolution year rather than inside it, close to McKinsey’s 30%-by-2030 (which is hours-weighted across all roles, not within-role), and still far more bullish than Forrester’s 6%-by-2030 (jobs-eliminated, a different metric). The crowd’s picture is one of steady structural erosion through 2035 rather than a step change in 2027–28 — which, after this refresh, is closer to where I land too.

Policy / regulation

The single most material near-term policy lever is the EU AI Act, and as of July 2026 the deferral is settled law rather than a proposal. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026: standalone Annex III high-risk obligations move from 2 August 2026 to 2 December 2027, and AI embedded in products already covered by EU product-safety law (Annex I) to 2 August 2028 [46]. Crucially, this is deferred, not cancelled, and the parts that bite soonest did not move — Article 50 transparency and AI-content-labelling duties, the GPAI provider obligations in force since August 2025, and the Article 5 prohibited-practices regime in force since February 2025 all held their original schedule. The omnibus also expands the AI Office’s investigatory and enforcement authority, including on-site inspection powers and the ability to secure binding commitments from providers. High-risk categories still include employment decisions, education, biometrics, critical infrastructure, migration/asylum and credit scoring — a meaningful share of “knowledge-work decisions” that touch consequential outputs — and agentic workflows must log events for risk identification, not just final outputs [26][46]. Functionally: a hard ceiling on what “autonomous without human review” can mean in EU regulated domains, now with a stronger enforcer and a firm 2027 date.

The California picture hardened from ethics guidance into statute. The State Bar’s May 2026 proposal wove AI duties into six existing Rules of Professional Conduct — Rule 1.1 competence requiring a lawyer to “independently review, verify, and exercise professional judgment regarding any output generated by the technology,” with Rules 5.1/5.3 supervisory obligations explicitly directed at agentic tools that “autonomously perform tasks or workflows without human prompting” [21]. Then on 31 August 2026 the legislature passed SB 574, which requires attorneys to disclose AI use in court filings, personally read and verify every cited legal source, correct hallucinated material before submission, refrain from entering confidential client information into certain generative systems, and not delegate the practice of law to AI at all — with the prohibition extended to arbitrators, who may not delegate decision-making or rely on undisclosed AI-generated information outside the record. Courts can sanction violators, and California sanctions above $1,000 trigger automatic State Bar investigation [47]. Agents can be 30% of a lawyer’s task throughput only if the rule allows it; in California it now explicitly does not, at statutory rather than advisory force. ABA’s task force in late 2025 declared AI “moved from experiment to infrastructure for the legal profession” [27], but their guidance remains oversight-heavy — and the direction of travel through 2026 was toward more mandated human verification, not less.

US executive orders under the current administration have been pro-deployment (rolled back the Biden EO’s reporting requirements early 2025), so federal regulatory drag is low in the US. The bigger US risk is state-level disclosure laws (California, New York, Colorado) and NLRB rulings on automation-driven layoffs that could slow integration.

Professional associations beyond the bar are slower to formalize — accounting (AICPA), engineering (NSPE), medicine (AMA) are at the “guidance” stage, not binding rules. McKinsey itself is deploying thousands of AI agents internally [17], which functions as a market signal that consulting (a flagship knowledge-work role) is being automated by its own incumbent practitioners — bullish for fast crossing of the 30% threshold in white-collar professional services.

Sub-gates (upstream)

The upstream dependencies that must be true for the gate to pass:

  1. METR 50%-reliability time horizon ≥ 1 week (40 hours) — P50: 2028 (slipped from 2027, Sept 2026). METR’s SOTA is 17.41h (Claude Mythos Preview, added 2026-05-08) and has not moved in four months; the point estimate already exceeds METR’s stated reliability ceiling of 16h and carries a 95% CI of 8.5–55h. Published doubling is 128.7 days from 2023 on (CI 104–158) and 187.8 days all-time-stitched — not the 89 days assumed in June. Naive extrapolation from 17.41h at 128.7-day doubling puts 40h at roughly September 2026; the honest statement is that the instrument can no longer resolve it, and METR chose to ship a different metric (expenditure horizon, July 2026) instead of a longer task suite [36][37].
  2. SWE-bench Pro ≥ 75% — P50: 2027 (not cleared; the June 2026 “cleared” call does not survive re-checking the cited source). Scale’s standardized public leaderboard, which runs every model through identical scaffolding, tops out at 61.5% (Muse Spark 1.1), with GPT-5.4 xHigh at 59.1% and Claude Opus 4.6 thinking at 51.9%; the frontier Anthropic models have not been added to it. The 77.8%/80.3% figures are vendor-reported on different harnesses — directionally real, but not the contamination-resistant apples-to-apples number this sub-gate is defined on [38].
  3. τ²-Bench policy-adherent score ≥ 90% — P50: 2026, cleared on the independent harness. Artificial Analysis has GLM-5.2 (max) and JT-35B-Flash at 99.1% and GLM-4.7-Flash at 98.8%; vendor-mixed aggregates using different user simulators and pass^k scoring still top out ~88%. The caveat is bigger than the gap: 45–48% of the remaining τ²-bench failures are silent false successes that no LLM judge detects above AUROC 0.65 [45]. The score clears the compliance bar; the monitoring problem it was meant to proxy does not.
  4. Agent cost-per-task < $1 for 10-minute knowledge-work unit — P50: 2027. Per-token prices fell ~67–80% year-on-year into Q1 2026 (roughly $18.40 → $6.07 per M tokens on a blended basis), but the frontier moved the other way: GPT-6 Astra lists at $10 input / $50 output per M tokens ($20/$100 in fast mode), and OpenAI explicitly asked the market to judge it on price-per-task rather than per-token — an acknowledgement that agentic loops multiply token counts faster than unit prices fall [31].
  5. Long-context 1M-token reliable recall — P50: 2026, effectively cleared. GPT-6 Astra ships a 1.05M-token window (922K max input, 128K output) and scores 100% on MRCR v2 8-needle at 256K–512K and 96.3% at 512K–1M, against Sol’s 91.5% and 73.8% [30][32]. What remains open is cross-session persistent memory — a system property, not a context-length property, and one the horizon-gap literature flags as where compression silently drops earlier constraints [48].
  6. Multi-agent orchestration in production — P50: 2027. Shipping and normalising — Anthropic’s 2026 agentic-coding report frames the year as the shift from single assistants to agent teams running autonomously for hours or days, citing a Rakuten run where Claude Code completed a seven-hour implementation in a 12.5M-line vLLM codebase at 99.9% numerical accuracy [39]. Adoption, not capability, is now the binding constraint: ~88% of enterprise agent pilots still never reach production [42].

Cross-gate dependencies

The 30%-knowledge-work gate has the following non-trivial relationships with the other 10 gates in this set:

Strongest dependencyai-tutor-k8-parity-20mo. Same underlying tech stack (long-horizon, tool-using, multi-turn LLM agents with retrieval). If a knowledge-work agent can handle 30% of a software engineer’s tasks, a K-8 tutor at parity is fundamentally a packaging and safety-tuning problem, not a capability problem. Relation: enables. Strength: strong. A 6-month lag would be typical — knowledge-work agents reach 30% threshold, K-8 tutors at parity follow within a model generation.

Medium correlationautonomous-freight-delivery. Both depend on long-horizon reliability and regulator acceptance of unsupervised autonomous action in consequential domains. The capability progress is mostly independent (one is mostly perception/control, one is mostly language/reasoning), but the deployment timing correlates because both run into the same compliance / liability infrastructure. Relation: correlates. Strength: medium.

Weak correlationrobotaxi-unit-economics-5-cities and humanoid-retail-20k. Different stacks technically, but both are autonomous action gates and the regulatory / labor-policy backlash applies broadly. If society broadly accepts “AI agents handling 30% of knowledge work,” it’s marginally easier to accept “robots stocking shelves” — but the binding constraints are very different. Relation: correlates. Strength: weak.

Substitutesconstruction-robot-40pct-labor. If construction labor automation accelerates, the political/economic pressure to slow knowledge-work automation may rise as a labor-market protection response — though more likely both proceed in parallel. Relation: weak.

Unrelatedcell-meat-beef-parity, residential-solar-storage-0.04, metals-bom-30pct, evtol-1k-trips-major-city, smr-first-oecd-deployment. These are physical-world cost-curve gates that don’t share a meaningful capability or policy bottleneck with knowledge-work agent autonomy.

Downstream impact essay

Labor (primary). The 30%-tasks-autonomous threshold passing in any one knowledge-work role triggers a phase change in that role’s labor economics within 24–36 months. Customer support has already crossed (Klarna, 67%) and the consequence has been: ~50% headcount reduction in supported teams, role redefinition toward “AI-assisted escalation specialist,” and dropping wages for the residual humans. Software engineering is mid-cross today: the Reddit signal of “4 engineers doing the work of 12” matches the Microsoft/Meta/Cloudflare/Cisco/Coinbase/Paypal/Fidelity wave of 2026-Q2 layoffs framed as “AI-driven efficiency.” If P50 = 2029 is right, by 2032 expect: (a) entry-level white-collar hiring at large firms dropping 40–60% from 2024 levels, (b) wage compression in the bottom-two quintiles of knowledge-work (back-office, junior analyst, paralegal, content moderator, basic legal review), (c) wage expansion in the top quintile for humans who can effectively orchestrate teams of agents. This is not “mass unemployment” — it’s “the bottom rung of the career ladder evaporating” while senior workers get more leverage. The CS enrollment drop in 2026 is the early demographic signal. Political response: UBI is dead (Altman quietly walked back), what’s likely instead is some mix of (a) reskilling tax credits, (b) sector-specific job guarantees, (c) “human in the loop” mandates in regulated industries.

Education (secondary). If by 2029 a typical software-dev / analyst / paralegal role is 30%+ automated, the K-12 → college pipeline reconfigures. The signal is already visible: CS enrollment drop, ME/EE rises. What kids actually need to learn changes: (a) judgment and verification — knowing when an agent’s output is wrong, why, and how to fix it; (b) agent orchestration — being good at directing a team of agents is the new “being good at managing people”; (c) deep domain knowledge — generic prompt-engineering gets commoditized fast, but knowing a domain well enough to ask the right questions becomes the durable skill; (d) physical / human-bound skills — trades, healthcare, hands-on creative, hospitality. The kids who get a generic information-economy education in 2026 will graduate into 2030–2034, exactly when the 30%-threshold-in-mainstream-roles is biting hardest. Curricula should bias toward depth-in-domain + AI-leveraged output, not breadth in legacy white-collar skills.

Travel (tertiary). If knowledge-work agents handle 30% of tasks autonomously, the marginal value of a knowledge worker being physically co-located with their team falls — they’re managing agents that operate 24/7 from anywhere, and the agent doesn’t care which time zone the human is in. This entrenches remote work and weakens RTO mandates economically — though RTO is now driven by real-estate sunk cost and managerial preference, not productivity. The labor-market consequence: location decisions for skilled remote workers become tax-arbitrage decisions (Israel → Portugal, US → Mexico City, NYC → Miami). For Tel Aviv specifically, the Israeli tech sector becomes more attractive to globally-distributed talent because (a) remote-work normalization, (b) Israeli engineers were already accustomed to working with US clients at distance, (c) lower COL than US. The 2nd-order effect on travel: business travel for routine knowledge-work coordination drops further; but high-stakes deal-making, conferences, and leadership offsites get more valuable because they’re the parts of work that agents can’t do.

Sources

  1. METR, Measuring AI Ability to Complete Long Tasks — original 2025 paper introducing the 50%-time-horizon metric and 7-month doubling rate; Claude 3.7 Sonnet at ~50 min. Accessed 2026-05-13.
  2. METR, Time Horizon 1.1 — January 2026 update: Claude Opus 4.5 at 320 min, GPT-5 at 214 min, doubling since 2024 of 89 days. Accessed 2026-05-13.
  3. LLM-Stats SWE-Bench Verified leaderboard — GPT-5.5 88.7%, Claude Opus 4.7 87.6% as of April 2026; OpenAI’s contamination warning. Accessed 2026-05-13.
  4. Scale AI SWE-Bench Pro public leaderboard — contamination-resistant variant; GPT-5.4 xHigh leads at 59.10%, Claude Opus 4.6 thinking at 51.9%. Accessed 2026-05-13.
  5. Coasty Blog OSWorld benchmark results 2026 — GPT-5.5 78.7%, Claude Opus 4.6 72.7%, Claude Sonnet 4.6 72.5%, human baseline ~72%. Accessed 2026-05-13.
  6. TechBuzz, Cursor Hits $2B ARR — $100M Jan 2025 → $2B Feb 2026; 70% of Fortune 1000 customers; 1M+ DAU. Accessed 2026-05-13.
  7. Cognition, Devin’s 2025 Performance Review — 67% PR merge rate (up from 34%); deployment at Goldman, Santander, Dell, Cisco; Infosys partnership for global deployment. Accessed 2026-05-13.
  8. SiliconANGLE, Cognition $25B valuation talks — ARR $1M Sep 2024 → $73M Jun 2025; product expanded to enterprise IDE + code review. Accessed 2026-05-13.
  9. VentureBeat, Anthropic $30B revenue run-rate — Claude Code $2.5B run-rate, >1,000 customers spending >$1M annually, 80x growth. Accessed 2026-05-13.
  10. Klarna press release, AI assistant handles two-thirds of customer service chats in its first month — 67% autonomous resolution, 2.3M chats in month one, $40M/yr saved, with later partial walkback to human-hybrid. Accessed 2026-05-13.
  11. Kwa et al., arXiv:2503.14499, Measuring AI Ability to Complete Long Software Tasks — methodology, 7-month doubling, 5-year extrapolation to month-long tasks. Accessed 2026-05-13.
  12. HAL Princeton GAIA leaderboard — Claude Sonnet 4.5 at 74.6% (scaffolded); human baseline ~92%; Anthropic sweeps top 6. Accessed 2026-05-13.
  13. Sierra, τ²-Bench leaderboard via Artificial Analysis — policy-adherent customer-service eval; frontier models cluster 50–65% as of April 2026. Accessed 2026-05-13.
  14. Anthropic, Project Vend Phase 2 — Claude running an actual vending shop; Phase 1 failed economically, Phase 2 improved with multi-agent architecture but “gap between capable and completely robust remains wide.” Accessed 2026-05-13.
  15. Sacra, Cursor / Replit / Lovable revenue tracking — Replit $150M ARR Sep 2025 → $1B target end-2026; Lovable $200M ARR Nov 2025 from $100M in 8 months. Accessed 2026-05-13.
  16. Anthropic Economic Index, March 2026 report — automation dominant in 1P API traffic; +14% computer/math tasks 6 months; customer service highest exposure; new “automated trading” and “sales outreach” categories 2x+. Accessed 2026-05-13.
  17. McKinsey Global Institute, Generative AI and the Future of Work in America — up to 30% of US hours automatable by 2030 with genAI; STEM jumps from 14% to 30%; 12M job switches. Accessed 2026-05-13.
  18. r/cscareerquestions top posts, May 2026 — representative post “4 engineers doing the work of 12”; CS enrollment drop; Microsoft/Cisco/Meta layoff wave framed as AI efficiency. Accessed 2026-05-13.
  19. r/LocalLLaMA, Kimi K2.6 is a legit Opus 4.7 replacement — open-weight catching up on ~85% of frontier tasks; commoditization of agent stack accelerating. Accessed 2026-05-13.
  20. r/Lawyertalk, Clients wore meta camera glasses to our consult then had AI analyze it — AI analysis pricing into client expectations for legal services. Accessed 2026-05-13.
  21. LawSites, California Bar Proposes Rule Requiring Lawyers to Verify Every AI Output — May 2026 ethics rule explicitly addressing agentic AI; verification mandatory. Accessed 2026-05-13.
  22. r/singularity top posts, April–May 2026 — Altman walks back UBI, half-marathon broken by robot, Figure AI 24x production scale; bullish-sentiment baseline. Accessed 2026-05-13.
  23. Metaculus, AI as a Competent Programmer Before 2030 — community implied resolution close to ~70–80% by 2030; closely related to upstream of this gate. Accessed 2026-05-13.
  24. Metaculus Labor Automation Forecasting Hub — collection of related forecasting questions on hours automated, jobs displaced; modal community estimate aligns with McKinsey 30%-by-2030. Accessed 2026-05-13.
  25. Manifold, Will AI cause the US Unemployment Rate to exceed 10% before 2030? — current price 15–20% probability; reflects market view that task automation ≠ mass unemployment. Accessed 2026-05-13.
  26. Trilateral Research, EU AI Act Compliance Timeline 2025–2027 — high-risk deadline pushed to Dec 2, 2027; agentic AI logging requirements. Accessed 2026-05-13.
  27. LawSites, ABA Task Force: AI Has Moved From Experiment to Infrastructure — late-2025 ABA report on legal AI institutional status. Accessed 2026-05-13.
  28. Scale AI SWE-bench Pro Leaderboard — contamination-resistant benchmark on private codebases; Claude Mythos Preview 77.8%, Claude Fable 5 80.3% as of June 2026, clearing the 75% sub-gate threshold ~1yr ahead of schedule. Accessed 2026-06-16.
  29. METR, Time Horizon 1.1 — January 2026 update — confirms ~14h time horizon for public frontier models (Feb 2026) with ~7-month doubling rate sustained; 40h crossing (sub-gate metr-time-horizon-1-week) on track for early 2027. Accessed 2026-06-16.
  30. VentureBeat, “Welcome to the AGI era”: OpenAI launches GPT-6 Astra — 2026-09-03 launch; ARC-AGI-3 98.6–99.9%, FrontierMath Tier 4 v2 97.6%, GPQA Diamond 96%, ExploitBench 100%, BenchCAD 95.9%, DeepSWE v1.1 74.1%, OSWorld 2.0 72.6% at ~40 min/task vs Sol’s 65.7% at ~75 min; $10/$50 per M tokens standard, $20/$100 fast; notes OpenAI omitted GDPval from launch materials and that NVIDIA’s AVO scaffold reached 100% ARC-AGI-3 on Claude Opus 5. Accessed 2026-09-07.
  31. TechRepublic, OpenAI launches GPT-6 Astra as Brockman says the “AGI era” has arrived — Brockman “we are now in the AGI era”; 0% unauthorized-scope breaches vs 48% for Sol; Pachocki caveat that “progress in intelligence does not guarantee progress in alignment”; Brockman argues pricing should shift from per-token to per-task. Accessed 2026-09-07.
  32. OpenAI developer docs, GPT-6 Astra model card — primary: 1,050,000-token context (922K max input, 128K output), $10/$1 cached/$50 per M tokens, computer use / web search / hosted shell via the Responses API. Cross-checked against Vellum’s benchmark table (AutomationBench 41.4% vs Fable 5.1’s 31.4% and Sol’s 18.1%; Terminal-Bench 4.0 57.7%; ScreenSpot-Pro 92.7%; MRCR v2 100% at 256K–512K). Accessed 2026-09-07.
  33. Artificial Analysis, GDPval-AA v2 leaderboard — 44 occupations across 9 sectors in agentic environments, blind pairwise LLM-judged Elo against a 1,000 human baseline: Claude Fable 5.1 (max) 1,766, Claude Opus 5 (max) 1,738, Muse Spark 1.3 1,720, GPT-5.6 Sol (max) 1,626 (18th), GPT-6 Astra (max) 1,582 (22nd). Accessed 2026-09-07.
  34. Artificial Analysis, AA-Briefcase — agentic knowledge-work benchmark — 91 tasks across data science, product management, banking operations and heavy-industry strategy, each scenario spanning up to six simulated weeks; Claude Fable 5.1 (max) 1,662 Elo leads, GPT-6 Astra (max) 6th at 1,562. Accessed 2026-09-07.
  35. The New Stack, OpenAI launches GPT-6 Astra and says welcome to the “AGI era” — independent write-up of the 2026-09-03 launch and benchmark table; Artificial Analysis Intelligence Index 61.2 for Astra vs 60.9 for Sol, behind Claude Fable 5.1 and Claude Opus 5 on OpenAI’s own comparison table. Accessed 2026-09-07.
  36. METR, Task-Completion Time Horizons of Frontier AI Models (live dashboard) — primary, read from the page’s embedded TH v1.1 dataset: doubling 128.744 days from 2023 on (CI 104.4–158.0) and 187.778 days all-time-stitched; Claude Opus 4.6 718.8 min, Claude Mythos Preview (early) 1,044.8 min (17.41h, CI 8.5–55.1h) still flagged is_sota; last dashboard update 2026-05-08; posted annotation “Measurements above 16 hrs are unreliable with our current task suite.” Accessed 2026-09-07.
  37. METR, Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT — July 2026 metric shipped in place of a longer time-horizon suite; GPT-5.5 ≈ $2.3K and Opus-4.8 ≈ $3.3K crossover against ~$2,500 of human labour per 1% improvement; “autonomous agent optimization has so far had minimal effect on AI R&D progress” on NanoGPT. Accessed 2026-09-07.
  38. Scale AI, SWE-Bench Pro public leaderboard (re-checked) — standardized identical-scaffolding public set as of Sept 2026: Muse Spark 1.1 61.50±3.10, GPT-5.4 xHigh 59.10±3.56, Muse Spark 55.00, Claude Opus 4.6 thinking 51.90, Gemini 3.1 Pro 46.10. No model on this leaderboard exceeds 75%; the 77.8/80.3% figures cited in the 2026-06-16 refresh are vendor-reported on different harnesses. Accessed 2026-09-07.
  39. Anthropic, 2026 Agentic Coding Trends Report — the tightest available read on this gate’s criterion: engineers use AI in roughly 60% of their work but can “fully delegate” only 0–20% of tasks; ~27% of AI-assisted work would not have been attempted at all; Rakuten case study of a seven-hour autonomous Claude Code implementation in a 12.5M-line vLLM codebase at 99.9% numerical accuracy. Accessed 2026-09-07.
  40. Anthropic, When AI builds itself — June 2026: >80% of code merged into Anthropic’s production codebase in May 2026 authored by Claude (from low single digits pre-Feb-2025); 8× code merged per engineer per day vs 2024; open-ended internal task success 26% (Nov 2025) → 76% (May 2026); research next-step judgment beating humans 51% → 64%. Explicit caveats: lines-of-code “almost certainly an overstatement,” self-reported 4× productivity likely high, automated review catches ~⅓ of bugs, and “human code review has become a new bottleneck.” Accessed 2026-09-07.
  41. ChatForest / Bloomberg reporting, Cognition ARR and valuation, September 2026 — $492M run-rate revenue May 2026 → >$900M by September, enterprise usage up >10× since the start of 2026, 89% of Cognition’s own committed code written by Devin, reported round at $47B (from $26B in May). Accessed 2026-09-07.
  42. Institute of Project Management, Why 88% of enterprise AI pilots never reach production — 88% of agent pilots fail to graduate to production; blockers cited as evaluation gaps 64%, governance friction 57%, model reliability 51%; March 2026 survey has 78% of enterprises running pilots with <15% reaching production; MIT NANDA’s review of 300+ disclosed deployments put 95% at zero measurable return. Accessed 2026-09-07.
  43. Anthropic Economic Index, Cadences (June 2026 report) — published 2026-06-26; measurable Claude usage on only 7.5% of 17,998 O*NET tasks; knowledge work clusters at the top of adoption with physical work near zero; respondent skew (Computer & Mathematical ~30% of respondents vs 4% of US employment); reported exposure rises with automation share. Accessed 2026-09-07.
  44. Challenger, Gray & Christmas, August 2026 job-cut report — 52,881 announced cuts in August (from 33,429 in July), restructuring first at 16,173 and AI fourth at 3,462; 529,914 cuts January–August 2026, down 41% year-on-year and the lowest Jan–Aug total since 2022. Accessed 2026-09-07.
  45. Advani, From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents, arXiv:2606.09863 — 1 Jun 2026, FAGEN@ICML2026; 9,876 τ²-bench trajectories (8 model families) and 1,879 AppWorld trajectories (4 families); false success = 45–48% of failures in single-control τ²-bench domains, 3% in dual-control telecom, 75.8% in AppWorld self-assessing coding agents; no LLM-judge configuration exceeds AUROC 0.65 on τ²-bench (0.54 on AppWorld); TF-IDF detectors reach 0.83/0.95 at 3,300× lower latency. Accessed 2026-09-07.
  46. Gibson Dunn, EU AI Act Omnibus Agreement — postponed high-risk deadlines and other key changes — Digital Omnibus on AI, Regulation (EU) 2026/1744, published in the Official Journal 2026-07-24 and in force 2026-07-27: Annex III standalone high-risk deferred to 2027-12-02, Annex I to 2028-08-02; Article 50 transparency, GPAI provider duties and Article 5 prohibitions unchanged; AI Office gains on-site inspection and binding-commitment powers. Accessed 2026-09-07.
  47. Hoodline, California lawmakers pass first-of-its-kind law on AI in court filings (SB 574) — legislature approved SB 574 on 2026-08-31: mandatory disclosure of AI use in court documents, personal verification of every cited legal source, correction of hallucinated material, restrictions on confidential-data entry, and a prohibition on delegating the practice of law to AI, extended to arbitrators; court sanctions above $1,000 trigger automatic State Bar investigation. Accessed 2026-09-07.
  48. Chen, Wang & Qu, The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents, arXiv:2608.06663 — 7 Aug 2026 survey of 1,547 papers (2024–2026); disambiguates long-horizon / long-context / long-term memory; finds outcome-only signals grow uninformative as horizons lengthen, with coherence collapse, goal drift, and compression that “silently erases the safety constraints” set at trajectory start; notes one leaderboard-topping SWE-bench system falling 12.47% → 3.97% after contamination and weak-test filtering. Accessed 2026-09-07.
  49. Srinivasan, Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations, arXiv:2608.11323 — 11 Aug 2026; across TheAgentCompany, τ²-bench and AppWorld the agent main effect explains <3% of total variance while agent-by-task interaction explains 7–23% (“leaderboards rank specialization, not capability”); reliability collapses on the hardest task quartile (Eρ² 0.752 → 0.000); training-cell reliability anti-correlates with held-out reliability (r = −0.90). Accessed 2026-09-07.
  50. MIT Technology Review, AI’s recursive self-improvement might not come so quickly after all — 18 Aug 2026; Princeton-led evaluation gave Claude Opus 4.8 agents six days and $3,000 to attack unpublished NeurIPS 2026 research questions: “capable of all the engineering required” but “unambiguously bad at carrying out the research itself”; both agent-generated papers rejected by the original authors grading them. Accessed 2026-09-07.
Full markdown source (frontmatter + body) ▾
---
title: AI agent autonomously handles 30%+ of a typical knowledge-work job
dimensions: ["labor","education","travel"]
horizon: medium
trigger: An AI agent autonomously completes ≥ 30% of tasks in a typical knowledge-work role (e.g. legal review / software dev / content / customer support) without human review on each step. 'Typical' = the modal task distribution of the role per O*NET or comparable taxonomy.
timeline: {"p10":2028,"p50":2029,"p90":2033}
confidence: medium
sub_gates: [{"slug":"metr-time-horizon-1-week","p50":2028,"why":"METR 50% time horizon crosses 40 working hours — agents can chain a typical week of work. Slipped from 2027 in Sept 2026: METR's dashboard SOTA has been static at 17.41h since Claude Mythos Preview was added on 2026-05-08, past METR's own posted ceiling ('measurements above 16 hrs are unreliable with our current task suite'), and the published doubling estimate is 128.7d (2023-on) / 187.8d (all-time stitched), not 89d. The threshold is currently unmeasurable on its stated instrument."},{"slug":"swe-bench-pro-75pct","p50":2027,"why":"Contamination-resistant SWE-bench Pro at 75%+ implies SWE agents handle the modal task autonomously. NOT cleared on re-check (Sept 2026): Scale's standardized public leaderboard — every model through identical scaffolding — tops out at 61.5% (Muse Spark 1.1), with GPT-5.4 xHigh at 59.1%. The 77.8/80.3% figures cited in June 2026 are vendor-reported on different harnesses, not the standardized set."},{"slug":"tau2-bench-policy-90pct","p50":2026,"why":"Customer-service agents hit 90% on policy-adherent multi-turn flows. Cleared on Artificial Analysis's independent harness (GLM-5.2 max and JT-35B-Flash both 99.1%, GLM-4.7-Flash 98.8%), though vendor-mixed aggregates still top out ~88% and arXiv:2606.09863 finds 45-48% of remaining τ²-bench failures are silent false successes that LLM judges cannot detect (AUROC ≤0.65) — the score clears the bar, the monitoring problem does not."},{"slug":"agent-cost-per-task-sub-1usd","p50":2027,"why":"Frontier agent task cost drops below $1 for a typical 10-minute knowledge-work unit, enabling per-task economics. Mixed: per-token prices fell ~67-80% year-on-year to Q1 2026, but frontier per-token pricing went the other way — GPT-6 Astra lists at $10/$50 per M tokens ($20/$100 fast mode) and OpenAI explicitly asked the market to judge it on price-per-task rather than per-token, because agentic loops multiply token counts."},{"slug":"long-context-1m-reliable","p50":2026,"why":"1M-token context with reliable recall across multi-day sessions. Effectively cleared Sept 2026: GPT-6 Astra ships a 1.05M-token window (922K max input) and scores 100% on MRCR v2 8-needle at 256K-512K and 96.3% at 512K-1M. Cross-session persistent memory, as distinct from in-context recall, remains the open half."},{"slug":"agent-orchestration-prod-maturity","p50":2027,"why":"Multi-agent orchestration (Claude Code subagents / Devin-style) ships GA at top-3 enterprise vendors. Shipping, but adoption is the constraint: ~88% of enterprise agent pilots still fail to reach production, with evaluation gaps (64%), governance friction (57%) and model reliability (51%) the cited blockers."}]
history: [{"date":"2026-09-07T00:00:00.000Z","p10":2028,"p50":2029,"p90":2033,"why":"Reverts the June pull-in: both legs of it failed re-verification, and the first direct measurement of the gate's own criterion landed below the bar. (1) METR — the dashboard's SOTA has been frozen at 17.41h (Claude Mythos Preview, released 2026-04-07, added 2026-05-08) with no new model added in four months, above METR's own posted ceiling 'Measurements above 16 hrs are unreliable with our current task suite'; the published doubling estimate is now 128.7d (2023-on, CI 104-158) and 187.8d all-time-stitched, not the 89d this gate assumed, and METR pivoted to a new 'expenditure horizon' metric in July 2026. The gate's own METR dependency prices a doubling-rate reversion at +1-2y; taking the low end gives +1y. (2) SWE-bench Pro — Scale's standardized public set (identical scaffolding) tops out at 61.5% (Muse Spark 1.1), not the 77.8/80.3% vendor-harness figures the June entry read off it, so the contamination-resistant 75% sub-gate is not cleared. (3) Anthropic's 2026 Agentic Coding Trends Report measures the criterion directly: engineers use AI in ~60% of their work but can 'fully delegate' only 0-20% of tasks; Anthropic's own 'When AI builds itself' reports >80% of merged code Claude-authored while naming human code review the new bottleneck — the 'without human review on each step' clause is exactly what has not happened. GPT-6 Astra (2026-09-03) is a real capability jump but ranks 22nd on GDPval-AA v2 (1,582 vs Fable 5.1's 1,766) and 6th on AA-Briefcase, and OpenAI omitted GDPval from launch materials. Confidence stays medium: capability and commercial momentum are intact, so P90 holds at 2033. — https://metr.org/time-horizons/; https://labs.scale.com/leaderboard/swe_bench_pro_public; https://resources.anthropic.com/2026-agentic-coding-trends-report; https://www.anthropic.com/institute/recursive-self-improvement; https://artificialanalysis.ai/evaluations/gdpval-aa"},{"date":"2026-06-16T00:00:00.000Z","p10":2027,"p50":2028,"p90":2033,"why":"SWE-bench Pro (Scale AI, contamination-resistant) cleared: Claude Mythos Preview 77.8% / Claude Fable 5 80.3% as of June 2026 — sub-gate swe-bench-pro-75pct resolves ~1yr ahead of schedule. METR 50% time horizon confirmed ~14h (public models, Feb 2026) with ~7-month doubling rate sustaining → 40h crossing (metr-time-horizon-1-week) on track for early 2027, ~1yr ahead of prior P50 2028. Both binding sub-gates shifting 1yr earlier cascades main P50 from 2029→2028. — https://labs.scale.com/leaderboard/swe_bench_pro_public; https://metr.org/blog/2026-1-29-time-horizon-1-1/"},{"date":"2026-05-13T00:00:00.000Z","p10":2027,"p50":2029,"p90":2034,"why":"Initial estimate from initial research."}]
cross_gate: [{"other":"humanoid-retail-20k","relation":"correlates","strength":"weak","note":"Both are 'automation gates' but in different domains; share macro labor-policy backlash but capability progress is largely independent."},{"other":"ai-tutor-k8-parity-20mo","relation":"enables","strength":"strong","note":"Same underlying agent stack — long-horizon, tool-using, multi-turn. If knowledge-work agents cross 30%, K-8 tutors at parity is essentially a packaging problem."},{"other":"construction-robot-40pct-labor","relation":"correlates","strength":"weak","note":"Embodiment is the binding constraint there, not cognition; agentic LLM progress helps planning layer only."},{"other":"autonomous-freight-delivery","relation":"correlates","strength":"medium","note":"Both depend on long-horizon reliability and regulatory acceptance of unsupervised autonomous action — common bottleneck."},{"other":"robotaxi-unit-economics-5-cities","relation":"correlates","strength":"weak","note":"Different perception/control stack but same regulator question: 'when do we let it act unsupervised?'"},{"other":"quantum-shor-2048bit","relation":"correlates","strength":"weak","note":"AI agents may accelerate quantum algorithm design and error-correction optimization (e.g., Google's AlphaQubit), but quantum computing's binding constraint is hardware engineering. Indirect connection."},{"other":"human-aging-halted","relation":"enables","strength":"strong","note":"AI agents handling 30%+ of knowledge work directly accelerates drug discovery — Retro+OpenAI made reprogramming 50x more efficient; Insilico's AI-discovered drug reached Phase IIa in under 4 years vs 8-12 conventionally. Compresses the aging gate's P10 by 3-5 years."},{"other":"corporate-sovereignty-territory","relation":"correlates","strength":"medium","note":"Remote work via AI agents decouples employment from location, increasing demand for alternative citizenship/residency products that corporate territories could supply."},{"other":"brain-in-vat-body-replacement","relation":"enables","strength":"medium","note":"AI accelerates immunosuppressant drug discovery, CRISPR design for pig xenografts, organ-on-chip simulation, surgical planning, closed-loop artificial organ control, and BCI signal decoding. Pulls body-replacement P10 forward by ~5 years without changing fundamental biological constraints."},{"other":"us-unemployment-10pct-12mo","relation":"enables","strength":"strong","note":"Most direct upstream gate for mass unemployment. 30%+ knowledge work substitution implies ~18% gross displacement of US labor force; net unemployment depends on reabsorption speed. Even Goldman base case adds 0.6pp; getting to sustained 10%+ requires AI substitution 3-5x faster than baseline."}]
key_dependencies: [{"factor":"METR doubling rate sustains or reverts","kind":"data","direction":"both","linked_gate":null,"impact":"PARTIALLY FLIPPED Sept 2026 and worth 1 of the priced 1-2 years. METR's published doubling estimate is now 128.7 days (2023-on, CI 104-158) and 187.8 days all-time-stitched, not the 89 days this gate assumed; the dashboard SOTA has been static at 17.41h (Claude Mythos Preview) since 2026-05-08 with no new model added. A full reversion to the ~7-month all-time rate from here would cost another ~1 year; a resumed sub-100-day rate confirmed on a re-scaled task suite would return P50 to 2028."},{"factor":"METR time-horizon instrument saturation","kind":"data","direction":"delays","linked_gate":null,"impact":"METR posts that 'measurements above 16 hrs are unreliable with our current task suite' and the SOTA point estimate (17.41h) already sits above it with a 95% CI of 8.5-55h; METR shipped a different metric ('expenditure horizon') in July 2026 rather than extending the suite. Until a longer-task suite ships, the 40h sub-gate cannot resolve either way, which caps confidence at medium regardless of underlying capability."},{"factor":"EU AI Act high-risk agentic ban","kind":"regulation","direction":"delays","linked_gate":null,"impact":"Deadline now settled in law, not proposed: the Digital Omnibus on AI (Regulation (EU) 2026/1744) entered into force 2026-07-27, deferring standalone Annex III high-risk obligations to 2027-12-02 (Annex I to 2028-08-02) while expanding AI Office inspection powers. A hard unsupervised-action reading at that deadline carves out 30-40% of EU knowledge-work roles and slows US adoption by precedent, dropping confidence one step."},{"factor":"Training data exhaustion 2027-28","kind":"capability","direction":"delays","linked_gate":null,"impact":"If synthetic data and RL-from-environment fail to compensate for data limits, the doubling rate breaks and P50 slides 3+ years to ~2032. Not yet flipped — frontier capability kept climbing through 2026 (ARC-AGI-3 and FrontierMath Tier 4 both effectively saturated by GPT-6 Astra) — but METR's July 2026 expenditure-horizon result (best agents ~$2-3K crossover vs ~$2,500 per 1% of human labor on NanoGPT, 'minimal effect on AI R&D progress') is the first quantitative sign that the self-acceleration channel is not yet compensating."},{"factor":"Enterprise integration friction vs capability gap","kind":"data","direction":"accelerates","linked_gate":null,"impact":"Resolving toward the slow branch: ~88% of enterprise agent pilots still fail to graduate to production (<15% of pilot-running enterprises reach it), blockers cited as evaluation gaps 64% / governance 57% / model reliability 51%, and the Anthropic Economic Index finds measurable Claude usage on only 7.5% of 17,998 O*NET tasks. The 'plumbing alone gets there' branch that would have opened P10 to 2027 is not the branch we are on; a reversal — pilot-to-production conversion above ~40% — would pull P50 back to 2028."},{"factor":"California Bar and sectoral AI oversight rules","kind":"regulation","direction":"delays","linked_gate":null,"impact":"Hardened from ethics proposal to statute: California SB 574 passed the legislature 2026-08-31, requiring attorneys to personally read and verify every AI-cited source, disclose AI use in court filings, and barring delegation of the practice of law to AI — extended to arbitrators. Together with the State Bar's May 2026 Rule 1.1 / 5.1 / 5.3 amendments covering agentic tools, mandatory human-verification in legal, finance and HR blocks the autonomous-without-review threshold in regulated roles, delaying sectoral P50 by 2+ years."},{"factor":"Unmonitored task-share measurement exists at all","kind":"capability","direction":"delays","linked_gate":null,"impact":"The gate resolves on a measured O*NET-referenced share of tasks done without per-step review, and no such measurement program exists. The closest direct read is Anthropic's 2026 Agentic Coding Trends Report — engineers use AI in ~60% of their work but can 'fully delegate' only 0-20% of tasks — i.e. below the 30% bar in the most AI-saturated role. A credible published role-level audit crossing 30% would move P50 in by 1-2 years on its own; continued absence of any instrument holds P50 at 2029 or later."},{"factor":"AI tutor K-8 parity shared stack","kind":"gate","direction":"accelerates","linked_gate":"ai-tutor-k8-parity-20mo","impact":"Shared long-horizon tool-using agent stack means crossing 30% in knowledge work essentially confirms K-8 parity is a packaging problem; the Sept 2026 slip of this gate's P50 to 2029 pushes the K-8 knock-on out by the same ~1 year, and a re-acceleration to 2028 would pull K-8 P50 in by ~6 months."},{"factor":"US mass unemployment trigger threshold","kind":"gate","direction":"accelerates","linked_gate":"us-unemployment-10pct-12mo","impact":"30%+ knowledge-work substitution is the most direct upstream driver of sustained 10% US unemployment; faster-than-P50 crossing compresses the unemployment gate's P10 by 2-3 years. Currently pointing the other way — Challenger counted 529,914 announced US job cuts Jan-Aug 2026, down 41% year-on-year and the lowest January-August total since 2022, with AI falling to the fourth-most-cited reason in August (3,462)."}]
external_calibration: {"metaculus":"https://www.metaculus.com/labor-hub/","manifold":"https://manifold.markets/ahalekelly/will-ai-cause-the-us-unemployment-r","expert_consensus":"McKinsey Global Institute (2023, reaffirmed 2025): up to 30% of US work hours automatable by 2030 with genAI. Metaculus Labor Automation Forecasting Hub: forecasters expect US employment to fall ~3% by 2035 against BLS's +3% projection, with the most AI-exposed occupations shrinking 17.2% and young-graduate unemployment rising 6%→12%. Manifold 'US unemployment >10% caused by AI before 2030' has drifted 15-20% → 25% (161 holders). Anthropic's 2026 Agentic Coding Trends Report is the tightest read on this gate's own criterion: ~60% of engineers' work uses AI, but only 0-20% of tasks can be fully delegated."}
last_updated: "2026-09-07T00:00:00.000Z"
sources_count: 50
---

## Why this refresh moved the timeline (September 2026)

GPT-6 Astra (2026-09-03) triggered this refresh, but it is not what moved it. Astra is a genuine capability jump — ARC-AGI-3 at 99.9%, FrontierMath Tier 4 at 97.6%, OSWorld 2.0 at 72.6% in ~40 min/task against Sol's 65.7% at ~75 min, a 1.05M-token window with 100% MRCR v2 at 256K-512K — and Greg Brockman used the launch to declare "we are now in the AGI era" [30][31][32]. It is also *mid-pack on exactly the work this gate is about*: 22nd on Artificial Analysis's GDPval-AA v2 (1,582 Elo at max effort against Claude Fable 5.1's 1,766) and 6th on AA-Briefcase's long-horizon knowledge-work suite, and OpenAI omitted GDPval — its own benchmark for economically valuable real-world tasks — from the launch materials entirely [33][34][35]. That is a frontier-model release, which this gate already prices in.

What moved the timeline is that **both legs of the June 2026 pull-in failed re-verification, and the first direct measurement of this gate's actual criterion came in below the bar.** (1) METR's dashboard SOTA has been frozen at **17.41 hours** since Claude Mythos Preview was added on 2026-05-08 — no new model in four months — and that point estimate already sits above METR's own posted warning that "measurements above 16 hrs are unreliable with our current task suite," with a 95% CI of 8.5-55h. The doubling estimate METR now publishes is **128.7 days** (2023-on, CI 104-158) and 187.8 days all-time-stitched, not the 89 days this gate assumed; in July 2026 METR shipped a different metric ("expenditure horizon") rather than extending the task suite [36][37]. (2) Scale's **standardized** SWE-bench Pro public leaderboard — every model through identical scaffolding — tops out at **61.5%** (Muse Spark 1.1), with GPT-5.4 xHigh at 59.1%; the 77.8%/80.3% figures the June entry read as clearing the 75% sub-gate are vendor-reported on different harnesses [38]. (3) Anthropic's 2026 Agentic Coding Trends Report measures the criterion head-on: engineers use AI in **~60% of their work but can "fully delegate" only 0-20% of tasks** [39]. Anthropic's own *When AI builds itself* reports >80% of merged production code authored by Claude — and names **human code review as the new bottleneck**, with automated review catching about a third of bugs that reached production [40]. The "without human review on each step" clause is precisely what has not happened at the most AI-saturated employer in the world.

The gate's own METR dependency prices a doubling-rate reversion at +1-2 years. Taking the low end — because this is partly instrument saturation rather than demonstrated capability stall — gives **+1 year: P50 2028 → 2029, P10 2027 → 2028**. P90 stays 2033 and confidence stays medium: capability and commercial momentum are entirely intact (Cognition went from $492M to >$900M ARR between May and September 2026 at a reported $47B valuation, Cursor is at $2B, Anthropic's internal open-ended-task success went 26% → 76% in six months), and none of the new evidence lengthens the tail.

## TL;DR

I put the **P50 at 2029** — pushed back from 2028 in September 2026 when the two sub-gate claims behind the June pull-in did not survive re-checking — that an AI agent will autonomously complete ≥30% of tasks in at least one typical knowledge-work role, measured against an O*NET-style task inventory and without per-step human review. The headline thesis has not changed: customer support is the closest to crossing (Klarna handled 67% of conversations end-to-end, though it re-hired humans for the hard tail and now describes the hybrid as the destination rather than a retreat), and software engineering is visibly mid-cross with >80% of Anthropic's merged code Claude-authored and Devin writing 89% of Cognition's own commits. What the 2026 evidence sharpened is *which* gap is binding. It is not raw capability. It is the difference between **assisted throughput** (large, growing, measurable) and **unreviewed task share** (0-20% by the only direct survey we have), plus **integration friction** — ~88% of enterprise agent pilots never reach production — and **policy and liability**, which hardened rather than softened in 2026. **P10 = 2028** (one role, a fast-moving enterprise, formally audited against O*NET); **P90 = 2033** (a real stall in agent reliability or a hard regulatory clamp pushes the threshold out a decade). The single most important quantitative driver is still the METR time horizon — but as of September 2026 that instrument has stopped resolving, which is itself the news.

## Current state (as of 2026-09-07)

The 30% threshold is still not crossed in any agreed-upon "typical knowledge-work job" measured against O*NET, and the September 2026 evidence tightened rather than loosened that judgment. Four hard numbers anchor where we are:

- **METR time horizon**: SOTA is **17.41 hours** (Claude Mythos Preview, released 2026-04-07, added to the dashboard 2026-05-08) with a 95% CI of 8.5-55h — above METR's own stated reliability ceiling of 16h, and unchanged for four months despite GPT-5.6 Sol, Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra all shipping since. Published doubling: **128.7 days** from 2023 on (CI 104-158), 187.8 days all-time-stitched. A June 2026 pre-deployment eval of GPT-5.6 Sol returned only ~11.3h under standard scoring [36][37].
- **Delegation share**: Anthropic's 2026 Agentic Coding Trends Report — engineers use AI in **~60%** of their work but can *fully delegate* only **0-20%** of tasks; ~27% of AI-assisted work is work that would not have been attempted at all [39]. Anthropic internally: **>80%** of merged production code Claude-authored as of May 2026 (from low single digits pre-Claude-Code), engineers merging 8× the code per day versus 2024, open-ended internal task success 26% (Nov 2025) → 76% (May 2026) — with the caveat, from Anthropic, that lines-of-code "almost certainly overstates the true productivity gain" and that human review is now the bottleneck [40].
- **Benchmarks**: Scale's standardized **SWE-bench Pro** tops at **61.5%** (Muse Spark 1.1) [38]. **GDPval-AA v2** (44 occupations, 9 sectors, agentic environments, LLM-judged pairwise Elo against a 1,000 human baseline): Claude Fable 5.1 1,766, Claude Opus 5 1,738, GPT-6 Astra 1,582 [33]. **AutomationBench** (Zapier; 657 cross-application business-workflow tasks across finance, HR, marketing, operations, sales, support): best score **41.4%** [32]. **τ²-bench** clears 90% on Artificial Analysis's harness (99.1%) but ~88% on vendor-mixed aggregates. **OSWorld 2.0**: 72.6% [30][31].
- **Deployment and labor market**: Cognition at **>$900M ARR** (from $492M in May 2026) with 89% of its own code written by Devin and a reported $47B round; Cursor at $2B ARR [41]. Against that: **~88% of enterprise agent pilots fail to reach production**, with evaluation gaps (64%), governance friction (57%) and model reliability (51%) as blockers [42]; the Anthropic Economic Index finds measurable Claude usage on only **7.5% of 17,998 O*NET tasks** [43]; and Challenger counted **529,914** announced US job cuts January-August 2026, **down 41% year-on-year** and the lowest Jan-Aug total since 2022, with AI falling to the fourth-most-cited reason in August (3,462 cuts) [44].

So as of September 2026: capability keeps climbing, deployment revenue keeps compounding, and the *unreviewed* share of a real role remains the thing nobody can yet show above 30%. The lagging variable is still **measurement** — and it got worse, not better, this quarter: METR's series stopped resolving, Scale's standardized SWE-bench Pro set stopped receiving frontier models, and OpenAI declared the AGI era while withholding its own economic-work benchmark.

## Key uncertainties

1. **Has the METR doubling rate actually slowed, or has the instrument simply saturated?** METR's published estimate moved from 89 days (2024-on, TH 1.1 blog) to 128.7 days (2023-on, current dashboard), and the SOTA point has been static at 17.41h for four months — but METR also says anything above 16h is unreliable on the current suite and has not added the four frontier models released since. Extrapolating 17.41h at 128.7-day doubling puts 40h at roughly *now*; we simply cannot confirm it. Whether a re-scaled suite shows the trend intact or broken is the single highest-value observation for this gate.
2. **What % of tasks in a "typical" job are within agent capability today but blocked by integration / context / data access?** The 2026 read is that plumbing and governance dominate: ~88% of pilots die before production and only 7.5% of O*NET tasks show measurable frontier-model usage at all. That is the slow branch. Pilot-to-production conversion above ~40% would flip it.
3. **Can silent failure be monitored without a human in the loop?** This is the load-bearing question for "without per-step review" and it got a real answer in 2026: 45-48% of τ²-bench failures and 75.8% of AppWorld self-assessing coding-agent failures are *false successes*, and no LLM-judge configuration exceeds AUROC 0.65 at detecting them [45]. Cheap TF-IDF detectors do far better (0.83 / 0.95), which suggests the problem is tractable — but until unsupervised deployment has a monitor it can trust, per-step review is the rational default regardless of capability.
4. **Does the EU AI Act's December 2027 high-risk deadline functionally ban unsupervised agentic action in regulated domains (legal, finance, HR)?** No longer speculative on timing: the Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force 2026-07-27 — deferring Annex III high-risk obligations to 2027-12-02 and Annex I to 2028-08-02, while expanding the AI Office's inspection and binding-commitment powers [46]. Deferred, not cancelled, and the enforcement side got stronger.
5. **Does training data run out in 2027-28 and stall scaling?** Not yet visible in capability — ARC-AGI-3 and FrontierMath Tier 4 were effectively saturated in 2026 — but METR's expenditure-horizon result is the first quantitative hint that the self-acceleration channel isn't compensating: on the NanoGPT speedrun the best agents' crossover sits at ~$2-3K against ~$2,500 of human labor per 1% improvement, and METR concludes autonomous agent optimization "has so far had minimal effect on AI R&D progress" there [37].
6. **Will labor pushback and sectoral rules make 30%-autonomous a no-go even if the tech works?** Moving from rhetorical to statutory: California SB 574 passed the legislature 2026-08-31, requiring attorneys to personally verify every AI-cited source, disclose AI use in filings, and *not delegate the practice of law to AI* — extended to arbitrators [47].

## Evidence synthesis

### Academic

The single most important academic anchor is METR's *Measuring AI Ability to Complete Long Software Tasks* (Kwa et al., arXiv:2503.14499) [11], which introduces the time-horizon metric: the duration of human-expert task that the model can complete with 50% reliability. METR's longitudinal data shows a 7-month doubling from 2019–2025 (~9 seconds for GPT-3 in 2020 → ~50 minutes for Claude 3.7 in early 2025). The January 2026 TH 1.1 blog reported doubling compressing to **~89 days** since 2024 [2], which is the figure this gate ran on through June 2026.

**That figure no longer matches what METR publishes.** As of September 2026 the live dashboard's own embedded data gives doubling of **128.7 days** from 2023 on (95% CI 104.4–158.0) and **187.8 days** all-time-stitched, and the model series under TH 1.1 reads: Claude Opus 4.5 293 min (Nov 2025) → GPT-5.2 352 min (Dec 2025) → Claude Opus 4.6 **719 min ≈ 12.0h** (Feb 2026, CI 5.3–60.6h) → Claude Mythos Preview (early) **1,045 min ≈ 17.41h** (released 2026-04-07, added 2026-05-08, CI 8.5–55.1h). Mythos is still flagged `is_sota` four months later: GPT-5.6 Sol, Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra have not been added. METR's own chart annotation reads "**Measurements above 16 hrs are unreliable with our current task suite**" — i.e. the SOTA point estimate sits above the stated reliability ceiling, and a June 2026 pre-deployment evaluation of GPT-5.6 Sol returned only ~11.3h under standard scoring, with METR noting heavy sensitivity to reward-hacking behaviour [36]. In July 2026 METR published a *different* metric rather than a longer suite: the **expenditure horizon**, the dollar budget at which humans become more cost-effective than agents on continuously-scored optimisation. On the NanoGPT speedrun, the best models (GPT-5.5, Opus-4.8) reach crossover at roughly **$2.3K and $3.3K** against an estimated **$2,500 of human labour per 1% improvement** — i.e. about parity — and METR's own summary is that "autonomous agent optimization has so far had minimal effect on AI R&D progress" there [37]. Naive extrapolation from 17.41h at 128.7-day doubling puts the 40h threshold at roughly now; the defensible statement is that the instrument stopped resolving before it got there.

On benchmark-specific results, the **SWE-bench family** (Princeton/OpenAI, agent code repair) and its harder variants (SWE-bench Pro [4], SWE-bench Multimodal) show the same exponential — but with explicit warnings that Verified saturation is partly contamination [3]. The **GAIA benchmark** (Meta AI Research) for general AI assistants shows frontier models at 74.6% (Claude Sonnet 4.5, scaffolded) vs human baseline of ~92% [12]. **OSWorld** for computer-use agents has now been crossed by GPT-5.5 (78.7% vs ~72% human) [5]. **τ²-Bench** (Sierra) measures customer-service agents on policy-adherent multi-turn flows — a much harder bar than "task complete" — and frontier models cluster in the 50–65% range as of April 2026 [13].

Citation-graph leaders in agent eval are Princeton (SWE-bench, HAL), Sierra (τ-bench), CMU+OSU (WebArena, OSWorld), and METR itself. The trajectory across all five benchmarks is consistent: rapid progress on narrow tasks (coding), slower on broad open-world tasks (GAIA, OSWorld), and the slowest on multi-turn-with-policy (τ-bench) — which is also the most predictive of real deployment success.

Anthropic's own **Project Vend** [14] is informative as a negative result: in mid-2025 Claude failed to profitably run a tiny vending business autonomously over a month; in Phase 2 (with multi-agent architecture) it improved meaningfully but the gap between "capable" and "completely robust" remained wide. This is good evidence that *agentic business operation* is harder than *agentic task completion*, and is why I don't think benchmark scores alone get us to a 30%-of-a-job claim.

**Mid-2026 update — the reliability literature caught up with the capability literature.** Three post-June papers bear directly on whether per-step human review can be removed, and all three point the same way. *From Confident Closing to Silent Failure* (Advani, arXiv:2606.09863, FAGEN@ICML2026) characterises **false success** — the agent asserting completion when the environment state says otherwise — across 9,876 τ²-bench trajectories from 8 model families and 1,879 AppWorld trajectories from 4: **45-48% of failures in single-control τ²-bench domains, 3% in dual-control telecom, 75.8% among AppWorld self-assessing coding agents**. No LLM-judge configuration (5 judges × 5 prompt strategies, full task specs) exceeds **AUROC 0.65**; judges key on "confident closing language" rather than verified state changes. Lightweight TF-IDF detectors reach 0.83/0.95 at 3,300× lower latency, so this is fixable — but today the default monitor for an unsupervised agent cannot see half its failures [45]. *The Horizon Gap* (Chen et al., arXiv:2608.06663, 7 Aug 2026) surveys 1,547 papers and names the pattern: **outcome-only signals grow uninformative as horizons lengthen**, with coherence collapse and goal drift — including compression mechanisms that "silently erase the safety constraints" given at trajectory start — and notes SWE-bench resolution for one leaderboard-topping system falling from 12.47% to 3.97% after contamination and weak-test filtering [48]. *Deployment Decision Reliability* (Srinivasan, arXiv:2608.11323, 11 Aug 2026) decomposes variance across TheAgentCompany, τ²-bench and AppWorld and finds the **agent main effect explains <3% of total variance** while agent-by-task interaction explains 7-23%: "leaderboards rank specialization, not capability." Aggregate reliability collapses on the hardest task quartile (Eρ² on τ² action_checks falls 0.752 → 0.000), and training-cell reliability *negatively* correlates with held-out reliability (r = −0.90) — the designs that look most reliable replicate worst [49].

That literature is the reason I read GDPval-AA Elo and τ²-bench scores as necessary but nowhere near sufficient. It is also consistent with the Princeton-led open-ended-research evaluation reported in August 2026, where Claude Opus 4.8 agents given six days and $3,000 to attack unpublished NeurIPS 2026 questions were "capable of all the engineering required" but "unambiguously bad at carrying out the research itself" — both agent-generated papers were rejected by the original authors grading them [50].

On benchmark supply: **TheAgentCompany** (CMU) remains the closest public analogue to this gate's framing — a simulated software company where the best agents autonomously complete ~30% of tasks (39.3% with partial credit) — and its authors' conclusion still stands: agents "are not close to automating every task encountered in a workspace, even on the subset presented," with social interaction, professional UIs, and tasks lacking public documentation the hardest. Two newer suites are more directly on-point and both leave large headroom: **AutomationBench** (Zapier; 657 tasks across 47 simulated SaaS tools in finance, HR, marketing, operations, sales and support, scored on final environment state with no LLM judge) has a best score of 41.4%, and **AA-Briefcase** (91 tasks across four professional scenarios spanning up to six simulated weeks, producing spreadsheets/decks/memos) is led by Claude Fable 5.1 with GPT-6 Astra sixth [32][34].

### Industry / market

The deployment numbers from 2025–2026 are the strongest single piece of evidence that we are mid-crossing, not approaching, this gate. Five anchor points:

1. **Anthropic** crossed **$30B annualized revenue** by April 2026, up from $9B at end-2025 and $1B at end-2024 — an 80x growth in 16 months [9]. Claude Code alone is at $2.5B run-rate, with enterprise representing >50% of Claude Code revenue. >1,000 customers now spend >$1M annually with Anthropic, vs ~12 two years earlier.
2. **Cursor** hit **$2B ARR** by February 2026 from $100M one year earlier — used by ~70% of the Fortune 1000 [6]. Latest reporting puts Cursor in talks at a $50B valuation.
3. **Cognition / Devin**: PR merge rate up to **67%** in the 2025 performance review (from 34% the prior year); deployment at thousands of companies including Goldman Sachs, Santander, Dell, Cisco, Palantir; Infosys partnership for global financial-services rollout being "one of the largest agentic deployments to date" [7]. ARR went from $1M (Sep 2024) to $73M (Jun 2025). January 2026 launched **Devin Review** for automated code review.
4. **Replit** at $150M ARR (Sep 2025), on track for $1B run-rate by end-2026 [15]. **Lovable** at $200M ARR (Nov 2025), 8M users, 100K+ projects/day [15]. Both target the "non-engineer builds an app" segment which is itself a knowledge-work automation play.
5. **Klarna**: in customer support specifically, the 67% autonomous resolution rate (Feb 2024 launch month) achieved a 40% drop in cost-per-transaction over 2 years and replaced ~700 agents [10]. Note the 2025 walkback: Klarna re-introduced humans for the 5% of complex/emotional/edge-case conversations. This is the **clearest existing example** of crossing the 30% threshold in a real role today — though customer support is the *easiest* knowledge-work role to automate.

The **Anthropic Economic Index (March 2026)** [16] is the best public dataset on actual usage patterns: API traffic (which is more agentic / automated) shows the share of "computer and mathematical" tasks growing **+14%** over 6 months, and customer service is flagged as the occupational category with highest automation exposure. New API automation categories emerging in late 2025 include "business sales and outreach automation" and "automated trading and market operations" — both >2x growth. Critically, Anthropic notes that **automation already dominates 1P API traffic**, meaning that for the workloads developers ship to customers, the median interaction is already minimal-human-loop.

McKinsey Global Institute's 2023 estimate — reaffirmed in 2025 — pegs **up to 30% of US work hours automatable by 2030** with genAI [17]. For STEM specifically, McKinsey raises the 2030 automation potential from 14% to 30% under genAI scenarios. Forrester's 2024 estimate is more conservative (6% of US jobs eliminated by 2030), reflecting the distinction between *task automation* and *job elimination* that explains why my gate is well-defined: 30% of tasks within a role is a much earlier milestone than 30% of jobs gone.

**Mid-2026 update — the market got much bigger while the delegated share stayed small.** Cognition went from $492M run-rate revenue in May 2026 to **>$900M by September**, with 89% of its own committed code written by Devin and a reported round valuing it at $47B (up from $26B in May); Cursor's $2B ARR now carries a reported $60B acquisition option [41]. Anthropic's *When AI builds itself* (June 2026) is the single richest internal dataset anyone has published: **>80% of code merged into Anthropic's production codebase in May 2026 was authored by Claude** (from low single digits before Claude Code launched in Feb 2025); the typical engineer merges **8× as much code per day** as in 2024; success on the hardest, least-specified internal coding tasks went **26% (Nov 2025) → 76% (May 2026)**; and on 129 selected research-detour moments Mythos Preview beat the human's next-step choice 64% of the time, up from 51% for Opus 4.5 [40]. Anthropic supplies its own deflators, which matter for this gate: lines-of-code "measures quantity over quality… almost certainly an overstatement," the 4× self-reported productivity median from a March 2026 survey of 130 staff is probably high (they cite METR's own work on developers overestimating AI uplift), and — decisively — "**human code review has become a new bottleneck**," with humans shifting from writing code to only reviewing it. Automated Claude review would have caught roughly a third of the bugs behind past production incidents; the other two thirds are why review persists.

Set against that, the deployment funnel is the constraint. Roughly **88% of enterprise agent pilots never graduate to production**, and a March 2026 survey found 78% of enterprises running agent pilots with under 15% reaching production; the cited blockers are evaluation gaps (64%), governance friction (57%) and model reliability (51%) [42]. MIT's NANDA review of 300+ disclosed deployments put 95% of enterprise genAI pilots at zero measurable return. The **Anthropic Economic Index** (June 2026, "Cadences") is the cleanest O*NET-referenced read: measurable Claude usage appears on only **7.5% of 17,998 O*NET tasks**, knowledge work clusters at the top of adoption with physical work near zero, and the respondent base is heavily skewed (Computer & Mathematical occupations are ~30% of respondents against a 4% share of US employment) [43]. And the labor-market signal cooled: Challenger counted **529,914** announced US job cuts January-August 2026, **down 41% year-on-year** and the lowest Jan-Aug total since 2022, with AI slipping to the **fourth**-most-cited reason in August at 3,462 cuts, behind restructuring [44]. Klarna, still the clearest role-level example, now describes its hybrid — AI handling the two-thirds it handles well, humans rebuilt for the emotional and ambiguous tail — as the working end state rather than a retreat [10].

### Public sentiment

r/cscareerquestions in May 2026 is dominated by layoff posts and AI-displacement narratives. The top post — "4 engineers now doing the job of 12 at my friend's company because AI agents handle the rest" — has 1,315 upvotes and 516 comments [18]. The framing in earnings calls Microsoft, Meta, Cloudflare, Cisco, Coinbase, Paypal is consistent: "AI-driven efficiencies" justifying headcount cuts. CS enrollment is dropping for the first time in 6 years while ME and EE rise 11%/14% [18]. A senior FAANG manager's hopeful counter-narrative ("the golden age is coming, you're not cooked") got 1,354 upvotes — but the top comments push back hard. Sentiment among working software engineers is **directly observable as bearish on entry-level survival, bullish on senior-with-AI productivity**. The "first rung of the ladder turns into a button, eventually there's no ladder" meme has 1,602 upvotes.

r/LocalLLaMA shows a different signal: builders are skeptical of frontier-only narratives. Qwen 3.6 27B replacing Opus 4.7 for "85% of tasks" [19] suggests open-weight models are catching up fast — meaning the agent stack is commoditizing, not concentrating. This is bullish for "30% of role" timelines because it implies cost-per-task is falling fast enough to make economically viable agent deployment practical even for low-margin roles.

r/Lawyertalk is the most informative for legal-as-a-knowledge-work-role. Top post: clients showed up to an estate-planning consult wearing Meta camera glasses, had AI analyze the meeting, then sent back an AI-generated critique recommending an offshore trust and undercutting the lawyer's fees [20]. The DOJ-brief-clearly-written-by-AI post had 1,266 upvotes. Lawyers are simultaneously contemptuous of AI legal output quality *and* aware it's pricing into client expectations. California Bar's May 2026 rule proposal requires verification of every AI output [21] — this is exactly the friction that delays "30% autonomous" in regulated knowledge work even when the capability is there.

r/singularity is, predictably, on the bullish extreme — the modal post in April–May 2026 is robot manufacturing acceleration, Atlas tricks, half-marathon records broken by robots, GPT-5.4 solving 60-year-old Erdős problems [22]. Sam Altman publicly walked back UBI advocacy ("I no longer believe in universal basic income as much as I once did") — bullish signal that even OpenAI's CEO no longer thinks the labor displacement will be smoothed by cash transfers, which is itself an admission of expected speed.

### Prediction markets

The directly relevant **Metaculus** question is *AI as a Competent Programmer Before 2030* [23] — community resolution implies a median probability close to **70–80%** of competent autonomous programming by end-2029 under reasonable interpretations. The associated Metaculus "AI 2027" tournament aggregates questions on automation of AI R&D, with community estimates that AI begins automating AI research by **2027** — directly upstream of my gate. The **Metaculus Labor Automation Forecasting Hub** [24] hosts multiple questions on hours automated; the modal community estimate is consistent with McKinsey's 30%-by-2030 number.

On **Manifold**, the cleanest question I found is *"Will AI cause the US Unemployment Rate to exceed 10% before 2030?"* [25] — priced at **25%** as of September 2026 across 161 holders, up from the 15–20% recorded in May. Note the strict resolution bar: AI must be *directly* responsible for at least 3pp of the increase. This still sits well below any "AI handles 30% of tasks" question because *task automation* doesn't translate 1:1 to *unemployment* — many displaced workers shift roles rather than become unemployed.

The **Metaculus Labor Automation Forecasting Hub** [24] is now the better calibration anchor than the single programmer question, and its numbers are strikingly structural rather than dramatic: forecasters expect US employment to **fall ~3% by 2035** where BLS projects +3% growth (a six-point gap they attribute to AI); the most AI-exposed occupations shrink **17.2%** while the least-exposed grow 4.6%; labour's share of national income drops 62.1% → 55.6%; long-term unemployment rises 0.99% → 6.14%; young-graduate unemployment doubles 6% → 12%; and weekly hours fall 38 → 34 while wages still grow. Software developers and financial specialists are the largest projected decliners.

My P50 of 2029 now sits **roughly on** the Metaculus AI-programmer-by-2030 implied resolution year rather than inside it, close to McKinsey's 30%-by-2030 (which is hours-weighted across all roles, not within-role), and still far more bullish than Forrester's 6%-by-2030 (jobs-eliminated, a different metric). The crowd's picture is one of steady structural erosion through 2035 rather than a step change in 2027–28 — which, after this refresh, is closer to where I land too.

### Policy / regulation

The single most material near-term policy lever is the **EU AI Act**, and as of July 2026 the deferral is settled law rather than a proposal. The **Digital Omnibus on AI, Regulation (EU) 2026/1744**, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026: standalone Annex III high-risk obligations move from 2 August 2026 to **2 December 2027**, and AI embedded in products already covered by EU product-safety law (Annex I) to 2 August 2028 [46]. Crucially, this is *deferred, not cancelled*, and the parts that bite soonest did not move — Article 50 transparency and AI-content-labelling duties, the GPAI provider obligations in force since August 2025, and the Article 5 prohibited-practices regime in force since February 2025 all held their original schedule. The omnibus also **expands the AI Office's investigatory and enforcement authority**, including on-site inspection powers and the ability to secure binding commitments from providers. High-risk categories still include employment decisions, education, biometrics, critical infrastructure, migration/asylum and credit scoring — a meaningful share of "knowledge-work decisions" that touch consequential outputs — and agentic workflows must log events for risk identification, not just final outputs [26][46]. Functionally: a hard ceiling on what "autonomous without human review" can mean in EU regulated domains, now with a stronger enforcer and a firm 2027 date.

The **California** picture hardened from ethics guidance into statute. The State Bar's May 2026 proposal wove AI duties into six existing Rules of Professional Conduct — Rule 1.1 competence requiring a lawyer to "independently review, verify, and exercise professional judgment regarding any output generated by the technology," with Rules 5.1/5.3 supervisory obligations explicitly directed at agentic tools that "autonomously perform tasks or workflows without human prompting" [21]. Then on **31 August 2026 the legislature passed SB 574**, which requires attorneys to disclose AI use in court filings, personally read and verify every cited legal source, correct hallucinated material before submission, refrain from entering confidential client information into certain generative systems, and **not delegate the practice of law to AI at all** — with the prohibition extended to arbitrators, who may not delegate decision-making or rely on undisclosed AI-generated information outside the record. Courts can sanction violators, and California sanctions above $1,000 trigger automatic State Bar investigation [47]. Agents can be 30% of a lawyer's *task throughput* only if the rule allows it; in California it now explicitly does not, at statutory rather than advisory force. ABA's task force in late 2025 declared AI "moved from experiment to infrastructure for the legal profession" [27], but their guidance remains oversight-heavy — and the direction of travel through 2026 was toward more mandated human verification, not less.

US executive orders under the current administration have been pro-deployment (rolled back the Biden EO's reporting requirements early 2025), so federal regulatory drag is low in the US. The bigger US risk is **state-level disclosure laws** (California, New York, Colorado) and NLRB rulings on automation-driven layoffs that could slow integration.

Professional associations beyond the bar are slower to formalize — accounting (AICPA), engineering (NSPE), medicine (AMA) are at the "guidance" stage, not binding rules. McKinsey itself is deploying thousands of AI agents internally [17], which functions as a market signal that consulting (a flagship knowledge-work role) is being automated by its own incumbent practitioners — bullish for fast crossing of the 30% threshold in white-collar professional services.

## Sub-gates (upstream)

The upstream dependencies that must be true for the gate to pass:

1. **METR 50%-reliability time horizon ≥ 1 week (40 hours)** — P50: **2028** (slipped from 2027, Sept 2026). METR's SOTA is 17.41h (Claude Mythos Preview, added 2026-05-08) and has not moved in four months; the point estimate already exceeds METR's stated reliability ceiling of 16h and carries a 95% CI of 8.5–55h. Published doubling is 128.7 days from 2023 on (CI 104–158) and 187.8 days all-time-stitched — not the 89 days assumed in June. Naive extrapolation from 17.41h at 128.7-day doubling puts 40h at roughly September 2026; the honest statement is that **the instrument can no longer resolve it**, and METR chose to ship a different metric (expenditure horizon, July 2026) instead of a longer task suite [36][37].
2. **SWE-bench Pro ≥ 75%** — P50: **2027** (not cleared; the June 2026 "cleared" call does not survive re-checking the cited source). Scale's standardized public leaderboard, which runs every model through identical scaffolding, tops out at **61.5%** (Muse Spark 1.1), with GPT-5.4 xHigh at 59.1% and Claude Opus 4.6 thinking at 51.9%; the frontier Anthropic models have not been added to it. The 77.8%/80.3% figures are vendor-reported on different harnesses — directionally real, but not the contamination-resistant apples-to-apples number this sub-gate is defined on [38].
3. **τ²-Bench policy-adherent score ≥ 90%** — P50: **2026, cleared on the independent harness**. Artificial Analysis has GLM-5.2 (max) and JT-35B-Flash at 99.1% and GLM-4.7-Flash at 98.8%; vendor-mixed aggregates using different user simulators and pass^k scoring still top out ~88%. The caveat is bigger than the gap: 45–48% of the *remaining* τ²-bench failures are silent false successes that no LLM judge detects above AUROC 0.65 [45]. The score clears the compliance bar; the monitoring problem it was meant to proxy does not.
4. **Agent cost-per-task < $1 for 10-minute knowledge-work unit** — P50: 2027. Per-token prices fell ~67–80% year-on-year into Q1 2026 (roughly $18.40 → $6.07 per M tokens on a blended basis), but the frontier moved the other way: GPT-6 Astra lists at **$10 input / $50 output per M tokens** ($20/$100 in fast mode), and OpenAI explicitly asked the market to judge it on price-per-task rather than per-token — an acknowledgement that agentic loops multiply token counts faster than unit prices fall [31].
5. **Long-context 1M-token reliable recall** — P50: **2026, effectively cleared**. GPT-6 Astra ships a 1.05M-token window (922K max input, 128K output) and scores 100% on MRCR v2 8-needle at 256K–512K and 96.3% at 512K–1M, against Sol's 91.5% and 73.8% [30][32]. What remains open is *cross-session persistent memory* — a system property, not a context-length property, and one the horizon-gap literature flags as where compression silently drops earlier constraints [48].
6. **Multi-agent orchestration in production** — P50: 2027. Shipping and normalising — Anthropic's 2026 agentic-coding report frames the year as the shift from single assistants to agent teams running autonomously for hours or days, citing a Rakuten run where Claude Code completed a seven-hour implementation in a 12.5M-line vLLM codebase at 99.9% numerical accuracy [39]. Adoption, not capability, is now the binding constraint: ~88% of enterprise agent pilots still never reach production [42].

## Cross-gate dependencies

The 30%-knowledge-work gate has the following non-trivial relationships with the other 10 gates in this set:

**Strongest dependency** — `ai-tutor-k8-parity-20mo`. Same underlying tech stack (long-horizon, tool-using, multi-turn LLM agents with retrieval). If a knowledge-work agent can handle 30% of a software engineer's tasks, a K-8 tutor at parity is fundamentally a packaging and safety-tuning problem, not a capability problem. **Relation: enables. Strength: strong.** A 6-month lag would be typical — knowledge-work agents reach 30% threshold, K-8 tutors at parity follow within a model generation.

**Medium correlation** — `autonomous-freight-delivery`. Both depend on long-horizon reliability and regulator acceptance of unsupervised autonomous action in consequential domains. The *capability* progress is mostly independent (one is mostly perception/control, one is mostly language/reasoning), but the *deployment timing* correlates because both run into the same compliance / liability infrastructure. **Relation: correlates. Strength: medium.**

**Weak correlation** — `robotaxi-unit-economics-5-cities` and `humanoid-retail-20k`. Different stacks technically, but both are *autonomous action gates* and the regulatory / labor-policy backlash applies broadly. If society broadly accepts "AI agents handling 30% of knowledge work," it's marginally easier to accept "robots stocking shelves" — but the binding constraints are very different. **Relation: correlates. Strength: weak.**

**Substitutes** — `construction-robot-40pct-labor`. If construction labor automation accelerates, the political/economic pressure to slow knowledge-work automation may rise as a labor-market protection response — though more likely both proceed in parallel. **Relation: weak.**

**Unrelated** — `cell-meat-beef-parity`, `residential-solar-storage-0.04`, `metals-bom-30pct`, `evtol-1k-trips-major-city`, `smr-first-oecd-deployment`. These are physical-world cost-curve gates that don't share a meaningful capability or policy bottleneck with knowledge-work agent autonomy.

## Downstream impact essay

**Labor (primary).** The 30%-tasks-autonomous threshold passing in any one knowledge-work role triggers a phase change in that role's labor economics within 24–36 months. Customer support has already crossed (Klarna, 67%) and the consequence has been: ~50% headcount reduction in supported teams, role redefinition toward "AI-assisted escalation specialist," and dropping wages for the residual humans. Software engineering is mid-cross today: the Reddit signal of "4 engineers doing the work of 12" matches the Microsoft/Meta/Cloudflare/Cisco/Coinbase/Paypal/Fidelity wave of 2026-Q2 layoffs framed as "AI-driven efficiency." If P50 = 2029 is right, by 2032 expect: (a) entry-level white-collar hiring at large firms dropping 40–60% from 2024 levels, (b) wage compression in the bottom-two quintiles of knowledge-work (back-office, junior analyst, paralegal, content moderator, basic legal review), (c) wage expansion in the top quintile for humans who can effectively orchestrate teams of agents. This is not "mass unemployment" — it's "the bottom rung of the career ladder evaporating" while senior workers get more leverage. The CS enrollment drop in 2026 is the early demographic signal. Political response: UBI is dead (Altman quietly walked back), what's likely instead is some mix of (a) reskilling tax credits, (b) sector-specific job guarantees, (c) "human in the loop" mandates in regulated industries.

**Education (secondary).** If by 2029 a typical software-dev / analyst / paralegal role is 30%+ automated, the K-12 → college pipeline reconfigures. The signal is already visible: CS enrollment drop, ME/EE rises. What kids actually need to learn changes: (a) **judgment and verification** — knowing when an agent's output is wrong, why, and how to fix it; (b) **agent orchestration** — being good at directing a team of agents is the new "being good at managing people"; (c) **deep domain knowledge** — generic prompt-engineering gets commoditized fast, but knowing a domain well enough to ask the right questions becomes the durable skill; (d) **physical / human-bound skills** — trades, healthcare, hands-on creative, hospitality. The kids who get a generic information-economy education in 2026 will graduate into 2030–2034, exactly when the 30%-threshold-in-mainstream-roles is biting hardest. Curricula should bias toward depth-in-domain + AI-leveraged output, not breadth in legacy white-collar skills.

**Travel (tertiary).** If knowledge-work agents handle 30% of tasks autonomously, the marginal value of a knowledge worker being physically co-located with their team falls — they're managing agents that operate 24/7 from anywhere, and the agent doesn't care which time zone the human is in. This *entrenches* remote work and weakens RTO mandates economically — though RTO is now driven by real-estate sunk cost and managerial preference, not productivity. The labor-market consequence: location decisions for skilled remote workers become tax-arbitrage decisions (Israel → Portugal, US → Mexico City, NYC → Miami). For Tel Aviv specifically, the Israeli tech sector becomes *more* attractive to globally-distributed talent because (a) remote-work normalization, (b) Israeli engineers were already accustomed to working with US clients at distance, (c) lower COL than US. The 2nd-order effect on travel: business travel for routine knowledge-work coordination drops further; but high-stakes deal-making, conferences, and leadership offsites get *more* valuable because they're the parts of work that agents can't do.

## Sources

1. [METR, *Measuring AI Ability to Complete Long Tasks*](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/) — original 2025 paper introducing the 50%-time-horizon metric and 7-month doubling rate; Claude 3.7 Sonnet at ~50 min. Accessed 2026-05-13.
2. [METR, *Time Horizon 1.1*](https://metr.org/blog/2026-1-29-time-horizon-1-1/) — January 2026 update: Claude Opus 4.5 at 320 min, GPT-5 at 214 min, doubling since 2024 of 89 days. Accessed 2026-05-13.
3. [LLM-Stats SWE-Bench Verified leaderboard](https://llm-stats.com/benchmarks/swe-bench-verified) — GPT-5.5 88.7%, Claude Opus 4.7 87.6% as of April 2026; OpenAI's contamination warning. Accessed 2026-05-13.
4. [Scale AI SWE-Bench Pro public leaderboard](https://labs.scale.com/leaderboard/swe_bench_pro_public) — contamination-resistant variant; GPT-5.4 xHigh leads at 59.10%, Claude Opus 4.6 thinking at 51.9%. Accessed 2026-05-13.
5. [Coasty Blog OSWorld benchmark results 2026](https://coasty.ai/blog/osworld-benchmark-results-2026-computer-use-ranked) — GPT-5.5 78.7%, Claude Opus 4.6 72.7%, Claude Sonnet 4.6 72.5%, human baseline ~72%. Accessed 2026-05-13.
6. [TechBuzz, Cursor Hits $2B ARR](https://www.techbuzz.ai/articles/cursor-hits-2b-arr-doubles-revenue-in-just-3-months) — $100M Jan 2025 → $2B Feb 2026; 70% of Fortune 1000 customers; 1M+ DAU. Accessed 2026-05-13.
7. [Cognition, *Devin's 2025 Performance Review*](https://cognition.ai/blog/devin-annual-performance-review-2025) — 67% PR merge rate (up from 34%); deployment at Goldman, Santander, Dell, Cisco; Infosys partnership for global deployment. Accessed 2026-05-13.
8. [SiliconANGLE, Cognition $25B valuation talks](https://siliconangle.com/2026/04/23/cognition-creator-ai-software-engineer-devin-talks-raise-hundreds-millions-25b-valuation/) — ARR $1M Sep 2024 → $73M Jun 2025; product expanded to enterprise IDE + code review. Accessed 2026-05-13.
9. [VentureBeat, Anthropic $30B revenue run-rate](https://venturebeat.com/technology/anthropic-says-it-hit-a-30-billion-revenue-run-rate-after-crazy-80x-growth) — Claude Code $2.5B run-rate, >1,000 customers spending >$1M annually, 80x growth. Accessed 2026-05-13.
10. [Klarna press release, *AI assistant handles two-thirds of customer service chats in its first month*](https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/) — 67% autonomous resolution, 2.3M chats in month one, $40M/yr saved, with later partial walkback to human-hybrid. Accessed 2026-05-13.
11. [Kwa et al., arXiv:2503.14499, *Measuring AI Ability to Complete Long Software Tasks*](https://arxiv.org/abs/2503.14499) — methodology, 7-month doubling, 5-year extrapolation to month-long tasks. Accessed 2026-05-13.
12. [HAL Princeton GAIA leaderboard](https://hal.cs.princeton.edu/gaia) — Claude Sonnet 4.5 at 74.6% (scaffolded); human baseline ~92%; Anthropic sweeps top 6. Accessed 2026-05-13.
13. [Sierra, τ²-Bench leaderboard via Artificial Analysis](https://artificialanalysis.ai/evaluations/tau2-bench) — policy-adherent customer-service eval; frontier models cluster 50–65% as of April 2026. Accessed 2026-05-13.
14. [Anthropic, *Project Vend Phase 2*](https://www.anthropic.com/research/project-vend-2) — Claude running an actual vending shop; Phase 1 failed economically, Phase 2 improved with multi-agent architecture but "gap between capable and completely robust remains wide." Accessed 2026-05-13.
15. [Sacra, Cursor / Replit / Lovable revenue tracking](https://sacra.com/c/cursor/) — Replit $150M ARR Sep 2025 → $1B target end-2026; Lovable $200M ARR Nov 2025 from $100M in 8 months. Accessed 2026-05-13.
16. [Anthropic Economic Index, March 2026 report](https://www.anthropic.com/research/economic-index-march-2026-report) — automation dominant in 1P API traffic; +14% computer/math tasks 6 months; customer service highest exposure; new "automated trading" and "sales outreach" categories 2x+. Accessed 2026-05-13.
17. [McKinsey Global Institute, *Generative AI and the Future of Work in America*](https://www.mckinsey.com/mgi/our-research/generative-ai-and-the-future-of-work-in-america) — up to 30% of US hours automatable by 2030 with genAI; STEM jumps from 14% to 30%; 12M job switches. Accessed 2026-05-13.
18. [r/cscareerquestions top posts, May 2026](https://www.reddit.com/r/cscareerquestions/comments/1tb026z/4_engineers_now_doing_the_job_of_12_at_my_friends/) — representative post "4 engineers doing the work of 12"; CS enrollment drop; Microsoft/Cisco/Meta layoff wave framed as AI efficiency. Accessed 2026-05-13.
19. [r/LocalLLaMA, *Kimi K2.6 is a legit Opus 4.7 replacement*](https://www.reddit.com/r/LocalLLaMA/comments/1sr8p49/kimi_k26_is_a_legit_opus_47_replacement/) — open-weight catching up on ~85% of frontier tasks; commoditization of agent stack accelerating. Accessed 2026-05-13.
20. [r/Lawyertalk, *Clients wore meta camera glasses to our consult then had AI analyze it*](https://www.reddit.com/r/Lawyertalk/comments/1t6ssgo/clients_wore_meta_camera_glasses_to_our_consult/) — AI analysis pricing into client expectations for legal services. Accessed 2026-05-13.
21. [LawSites, *California Bar Proposes Rule Requiring Lawyers to Verify Every AI Output*](https://www.lawnext.com/2026/05/california-bar-proposes-rule-requiring-lawyers-to-verify-every-ai-output-and-five-other-ai-focused-ethics-changes.html) — May 2026 ethics rule explicitly addressing agentic AI; verification mandatory. Accessed 2026-05-13.
22. [r/singularity top posts, April–May 2026](https://www.reddit.com/r/singularity/comments/1t14fpg/sam_altman_no_longer_believes_in_universal_basic/) — Altman walks back UBI, half-marathon broken by robot, Figure AI 24x production scale; bullish-sentiment baseline. Accessed 2026-05-13.
23. [Metaculus, *AI as a Competent Programmer Before 2030*](https://www.metaculus.com/questions/11188/ai-as-a-competent-programmer-before-2030/) — community implied resolution close to ~70–80% by 2030; closely related to upstream of this gate. Accessed 2026-05-13.
24. [Metaculus Labor Automation Forecasting Hub](https://www.metaculus.com/labor-hub/) — collection of related forecasting questions on hours automated, jobs displaced; modal community estimate aligns with McKinsey 30%-by-2030. Accessed 2026-05-13.
25. [Manifold, *Will AI cause the US Unemployment Rate to exceed 10% before 2030?*](https://manifold.markets/ahalekelly/will-ai-cause-the-us-unemployment-r) — current price 15–20% probability; reflects market view that task automation ≠ mass unemployment. Accessed 2026-05-13.
26. [Trilateral Research, *EU AI Act Compliance Timeline 2025–2027*](https://trilateralresearch.com/responsible-ai/eu-ai-act-implementation-timeline-mapping-your-models-to-the-new-risk-tiers) — high-risk deadline pushed to Dec 2, 2027; agentic AI logging requirements. Accessed 2026-05-13.
27. [LawSites, *ABA Task Force: AI Has Moved From Experiment to Infrastructure*](https://www.lawnext.com/2025/12/aba-task-force-ai-has-moved-from-experiment-to-infrastructure-for-the-legal-profession.html) — late-2025 ABA report on legal AI institutional status. Accessed 2026-05-13.
28. [Scale AI SWE-bench Pro Leaderboard](https://labs.scale.com/leaderboard/swe_bench_pro_public) — contamination-resistant benchmark on private codebases; Claude Mythos Preview 77.8%, Claude Fable 5 80.3% as of June 2026, clearing the 75% sub-gate threshold ~1yr ahead of schedule. Accessed 2026-06-16.
29. [METR, *Time Horizon 1.1 — January 2026 update*](https://metr.org/blog/2026-1-29-time-horizon-1-1/) — confirms ~14h time horizon for public frontier models (Feb 2026) with ~7-month doubling rate sustained; 40h crossing (sub-gate metr-time-horizon-1-week) on track for early 2027. Accessed 2026-06-16.
30. [VentureBeat, *"Welcome to the AGI era": OpenAI launches GPT-6 Astra*](https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra) — 2026-09-03 launch; ARC-AGI-3 98.6–99.9%, FrontierMath Tier 4 v2 97.6%, GPQA Diamond 96%, ExploitBench 100%, BenchCAD 95.9%, DeepSWE v1.1 74.1%, OSWorld 2.0 72.6% at ~40 min/task vs Sol's 65.7% at ~75 min; $10/$50 per M tokens standard, $20/$100 fast; notes OpenAI omitted GDPval from launch materials and that NVIDIA's AVO scaffold reached 100% ARC-AGI-3 on Claude Opus 5. Accessed 2026-09-07.
31. [TechRepublic, *OpenAI launches GPT-6 Astra as Brockman says the "AGI era" has arrived*](https://www.techrepublic.com/article/news-openai-gpt-6-astra-agi-era-2026/) — Brockman "we are now in the AGI era"; 0% unauthorized-scope breaches vs 48% for Sol; Pachocki caveat that "progress in intelligence does not guarantee progress in alignment"; Brockman argues pricing should shift from per-token to per-task. Accessed 2026-09-07.
32. [OpenAI developer docs, *GPT-6 Astra model card*](https://developers.openai.com/api/docs/models/gpt-6-astra) — primary: 1,050,000-token context (922K max input, 128K output), $10/$1 cached/$50 per M tokens, computer use / web search / hosted shell via the Responses API. Cross-checked against Vellum's benchmark table (AutomationBench 41.4% vs Fable 5.1's 31.4% and Sol's 18.1%; Terminal-Bench 4.0 57.7%; ScreenSpot-Pro 92.7%; MRCR v2 100% at 256K–512K). Accessed 2026-09-07.
33. [Artificial Analysis, *GDPval-AA v2 leaderboard*](https://artificialanalysis.ai/evaluations/gdpval-aa) — 44 occupations across 9 sectors in agentic environments, blind pairwise LLM-judged Elo against a 1,000 human baseline: Claude Fable 5.1 (max) 1,766, Claude Opus 5 (max) 1,738, Muse Spark 1.3 1,720, GPT-5.6 Sol (max) 1,626 (18th), GPT-6 Astra (max) 1,582 (22nd). Accessed 2026-09-07.
34. [Artificial Analysis, *AA-Briefcase — agentic knowledge-work benchmark*](https://artificialanalysis.ai/evaluations/aa-briefcase) — 91 tasks across data science, product management, banking operations and heavy-industry strategy, each scenario spanning up to six simulated weeks; Claude Fable 5.1 (max) 1,662 Elo leads, GPT-6 Astra (max) 6th at 1,562. Accessed 2026-09-07.
35. [The New Stack, *OpenAI launches GPT-6 Astra and says welcome to the "AGI era"*](https://thenewstack.io/openai-gpt6-astra-benchmarks/) — independent write-up of the 2026-09-03 launch and benchmark table; Artificial Analysis Intelligence Index 61.2 for Astra vs 60.9 for Sol, behind Claude Fable 5.1 and Claude Opus 5 on OpenAI's own comparison table. Accessed 2026-09-07.
36. [METR, *Task-Completion Time Horizons of Frontier AI Models* (live dashboard)](https://metr.org/time-horizons/) — primary, read from the page's embedded TH v1.1 dataset: doubling 128.744 days from 2023 on (CI 104.4–158.0) and 187.778 days all-time-stitched; Claude Opus 4.6 718.8 min, Claude Mythos Preview (early) 1,044.8 min (17.41h, CI 8.5–55.1h) still flagged `is_sota`; last dashboard update 2026-05-08; posted annotation "Measurements above 16 hrs are unreliable with our current task suite." Accessed 2026-09-07.
37. [METR, *Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT*](https://metr.org/blog/2026-07-21-expenditure-horizon/) — July 2026 metric shipped in place of a longer time-horizon suite; GPT-5.5 ≈ $2.3K and Opus-4.8 ≈ $3.3K crossover against ~$2,500 of human labour per 1% improvement; "autonomous agent optimization has so far had minimal effect on AI R&D progress" on NanoGPT. Accessed 2026-09-07.
38. [Scale AI, *SWE-Bench Pro public leaderboard* (re-checked)](https://labs.scale.com/leaderboard/swe_bench_pro_public) — standardized identical-scaffolding public set as of Sept 2026: Muse Spark 1.1 61.50±3.10, GPT-5.4 xHigh 59.10±3.56, Muse Spark 55.00, Claude Opus 4.6 thinking 51.90, Gemini 3.1 Pro 46.10. No model on this leaderboard exceeds 75%; the 77.8/80.3% figures cited in the 2026-06-16 refresh are vendor-reported on different harnesses. Accessed 2026-09-07.
39. [Anthropic, *2026 Agentic Coding Trends Report*](https://resources.anthropic.com/2026-agentic-coding-trends-report) — the tightest available read on this gate's criterion: engineers use AI in roughly 60% of their work but can "fully delegate" only 0–20% of tasks; ~27% of AI-assisted work would not have been attempted at all; Rakuten case study of a seven-hour autonomous Claude Code implementation in a 12.5M-line vLLM codebase at 99.9% numerical accuracy. Accessed 2026-09-07.
40. [Anthropic, *When AI builds itself*](https://www.anthropic.com/institute/recursive-self-improvement) — June 2026: >80% of code merged into Anthropic's production codebase in May 2026 authored by Claude (from low single digits pre-Feb-2025); 8× code merged per engineer per day vs 2024; open-ended internal task success 26% (Nov 2025) → 76% (May 2026); research next-step judgment beating humans 51% → 64%. Explicit caveats: lines-of-code "almost certainly an overstatement," self-reported 4× productivity likely high, automated review catches ~⅓ of bugs, and "human code review has become a new bottleneck." Accessed 2026-09-07.
41. [ChatForest / Bloomberg reporting, *Cognition ARR and valuation, September 2026*](https://chatforest.com/reviews/cognition-ai-devin-1-billion-round-26b-valuation-492m-arr-2026/) — $492M run-rate revenue May 2026 → >$900M by September, enterprise usage up >10× since the start of 2026, 89% of Cognition's own committed code written by Devin, reported round at $47B (from $26B in May). Accessed 2026-09-07.
42. [Institute of Project Management, *Why 88% of enterprise AI pilots never reach production*](https://www.institutepm.com/knowledge-hub/why-enterprise-ai-pilots-fail) — 88% of agent pilots fail to graduate to production; blockers cited as evaluation gaps 64%, governance friction 57%, model reliability 51%; March 2026 survey has 78% of enterprises running pilots with <15% reaching production; MIT NANDA's review of 300+ disclosed deployments put 95% at zero measurable return. Accessed 2026-09-07.
43. [Anthropic Economic Index, *Cadences* (June 2026 report)](https://www.anthropic.com/research/economic-index-june-2026-report) — published 2026-06-26; measurable Claude usage on only 7.5% of 17,998 O*NET tasks; knowledge work clusters at the top of adoption with physical work near zero; respondent skew (Computer & Mathematical ~30% of respondents vs 4% of US employment); reported exposure rises with automation share. Accessed 2026-09-07.
44. [Challenger, Gray & Christmas, *August 2026 job-cut report*](https://tradingeconomics.com/united-states/challenger-job-cuts) — 52,881 announced cuts in August (from 33,429 in July), restructuring first at 16,173 and AI fourth at 3,462; 529,914 cuts January–August 2026, down 41% year-on-year and the lowest Jan–Aug total since 2022. Accessed 2026-09-07.
45. [Advani, *From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents*, arXiv:2606.09863](https://arxiv.org/abs/2606.09863) — 1 Jun 2026, FAGEN@ICML2026; 9,876 τ²-bench trajectories (8 model families) and 1,879 AppWorld trajectories (4 families); false success = 45–48% of failures in single-control τ²-bench domains, 3% in dual-control telecom, 75.8% in AppWorld self-assessing coding agents; no LLM-judge configuration exceeds AUROC 0.65 on τ²-bench (0.54 on AppWorld); TF-IDF detectors reach 0.83/0.95 at 3,300× lower latency. Accessed 2026-09-07.
46. [Gibson Dunn, *EU AI Act Omnibus Agreement — postponed high-risk deadlines and other key changes*](https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/) — Digital Omnibus on AI, Regulation (EU) 2026/1744, published in the Official Journal 2026-07-24 and in force 2026-07-27: Annex III standalone high-risk deferred to 2027-12-02, Annex I to 2028-08-02; Article 50 transparency, GPAI provider duties and Article 5 prohibitions unchanged; AI Office gains on-site inspection and binding-commitment powers. Accessed 2026-09-07.
47. [Hoodline, *California lawmakers pass first-of-its-kind law on AI in court filings (SB 574)*](https://hoodline.com/2026/09/california-lawmakers-pass-first-of-its-kind-law-cracking-down-on-ai-fibs-in-court/) — legislature approved SB 574 on 2026-08-31: mandatory disclosure of AI use in court documents, personal verification of every cited legal source, correction of hallucinated material, restrictions on confidential-data entry, and a prohibition on delegating the practice of law to AI, extended to arbitrators; court sanctions above $1,000 trigger automatic State Bar investigation. Accessed 2026-09-07.
48. [Chen, Wang & Qu, *The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents*, arXiv:2608.06663](https://arxiv.org/abs/2608.06663) — 7 Aug 2026 survey of 1,547 papers (2024–2026); disambiguates long-horizon / long-context / long-term memory; finds outcome-only signals grow uninformative as horizons lengthen, with coherence collapse, goal drift, and compression that "silently erases the safety constraints" set at trajectory start; notes one leaderboard-topping SWE-bench system falling 12.47% → 3.97% after contamination and weak-test filtering. Accessed 2026-09-07.
49. [Srinivasan, *Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations*, arXiv:2608.11323](https://arxiv.org/abs/2608.11323) — 11 Aug 2026; across TheAgentCompany, τ²-bench and AppWorld the agent main effect explains <3% of total variance while agent-by-task interaction explains 7–23% ("leaderboards rank specialization, not capability"); reliability collapses on the hardest task quartile (Eρ² 0.752 → 0.000); training-cell reliability anti-correlates with held-out reliability (r = −0.90). Accessed 2026-09-07.
50. [MIT Technology Review, *AI's recursive self-improvement might not come so quickly after all*](https://www.technologyreview.com/2026/08/18/1142188/ai-recursive-self-improvement/) — 18 Aug 2026; Princeton-led evaluation gave Claude Opus 4.8 agents six days and $3,000 to attack unpublished NeurIPS 2026 research questions: "capable of all the engineering required" but "unambiguously bad at carrying out the research itself"; both agent-generated papers rejected by the original authors grading them. Accessed 2026-09-07.