Recursive Self-Improvement at OpenAI and Anthropic · Section 7 of 9

7. Verdict, Base Rates, and What Would Change It

Version 1.20, revised 28 September 2026

7.1 The rung-by-rung verdict

Rung Verdict Basis
1. AI-assisted coding at labs Verified and near-saturated on adoption; the productivity multiplier is unverified >80% of merged code at Anthropic authored by Claude as of May 2026, with Anthropic's own caveat that lines of code overstate the true gain [Anthropic Institute, 2026]; AI embedded in day-to-day R&D at OpenAI and Google [METR, 2026b]. The only RCT found a 19% slowdown for experienced developers on early-2025 tools while they believed they were 20% faster [METR, 2025a], and the follow-up experiment was redesigned over selection effects [METR, 2026d]. [confidence: high on adoption; low-medium on the size of the gain]
2. Autonomous multi-hour engineering Verified in the 1–16 hour range at 50% reliability, with a degrading instrument 80% reliability sits roughly an order of magnitude lower (70 minutes against a 719-minute 50% horizon for Opus 4.6); METR states measurements above 16 hours are unreliable with the current suite, and detected cheating contaminated at least 16% of successful 8-hour-plus runs [METR, 2026a; METR, 2026b; METR, 2026f]. [confidence: high]
3. Autonomous novel research yielding real gains Verified narrowly where exact verifiers exist; falsified for open-ended research Kernels, scheduling heuristics, matrix multiplication, config adaptation [Google DeepMind, 2025/2026; OpenAI, 2026b]. Frontier agents given six days and ~$3,000 of compute completed the engineering of two unpublished NeurIPS 2026 submissions and failed the research; both papers were rejected by the original authors [Kirgis et al., 2026]. [confidence: high — labs and independent evaluators agree on the split]
4. Closed loop shortening the next cycle Not demonstrated, not claimed No model rated High in OpenAI's AI Self-Improvement category [OpenAI, 2026a]; METR concluded GPT-5.6 Sol "would not enable fully automated AI R&D" [METR, 2026c]; Anthropic's automated AI R&D threshold has not been declared crossed [Anthropic, 2026a]. [confidence: high]
5. Sustained superexponential, humans out of the loop Not demonstrated, not claimed by anyone No lab claims it; the one prediction market with explicit resolution criteria prices "RSI by mid-2026" at 5% (re-pulled August 23, 2026; volatility note in Section 3.5) [Manifold, 2026]. [confidence: high]

7.2 The structure of the expert disagreement

Severin Field's interviews with 25 researchers across OpenAI, Anthropic, Google DeepMind, Meta, Princeton, UC Berkeley and Stanford (conducted August–September 2025, published August 13, 2026) are the best available map [Field, 2026]. Three numbers carry the structure: 20 of 25 ranked automating AI R&D among the most severe and urgent risks from AI systems; of 21 who addressed trajectories, 12 expected scaling trends to continue until AI matches human researcher labor; 16 expressed skepticism about positive feedback loops specifically. Experts broadly expect capability parity and doubt the loop, and that split maps exactly onto the rung 3/4 distinction. The skeptics argue that paradigm-shifting breakthroughs require memory, creativity and genuine novelty that scaling has not delivered; the believers point to the METR horizon trend. (Field also found 17 of 25 expressing reservations about capable models being kept internal [Field, 2026].)

Field identifies three factors explaining why company researchers are more bullish than academics: selection effects (believers gravitate to well-funded labs), proximity to progress, and hype incentives [Field, 2026]. All three apply whenever a lab statement is read.

7.3 The forecast spread as of August 2026

Source Forecast Date Sourcing
Jack Clark (Anthropic) 60% chance of RSI by end-2028; ~30% as early as 2027 (the 2027 figure secondary) May 4, 2026 Primary for the 60%: X post [Clark, 2026, https://x.com/jackclarkSF/status/2051312759594471886]; Axios interview May 7, 2026
Anthropic Frontier Safety Roadmap Research automation "plausible, as soon as early 2027" (full statement, Section 2.2) July 10, 2026 Primary [Anthropic, 2026b]; a safeguards-planning document, not a prediction
OpenAI (Altman/Pachocki) Research-intern goal September 2026; automated researcher 2028 Oct 28, 2025 Livestream via TechCrunch [TechCrunch, 2025]
Jared Kaplan (Anthropic) Humanity decides "between 2027 and 2030" whether to let AI train itself Dec 2, 2025 Guardian interview; the circulating "as little as a year away" line has no primary source
Kokotajlo Superhuman-coder median ~end-2029/early-2030; AGI (TED-AI) median Dec 2030 Jan 27, 2026 Primary [Lifland, Kokotajlo & Halstead, 2026]
Lifland TED-AI median 2035 Jan 27, 2026 Primary [same]
Cotra 10% full AI R&D automation by end-2026; 24h METR horizon median Jan 14, 2026 Primary [Cotra, 2026]; by March 5, 2026 she said her SWE forecasts already "felt much too conservative"
METR pilot — AI experts 20% that six years of progress compresses into two Aug 2025 [METR, 2025d]; pilot, "suggestive evidence only" [confidence: medium — pilot figures not independently reconfirmed]
METR pilot — superforecasters 8% for the same question Aug 2025 [METR, 2025d]
Manifold market 5% RSI by mid-2026 (re-pulled August 23, 2026; volatility note in Section 3.5) Aug 2026 [Manifold, 2026]
Metaculus community 25% AGI by 2029; 50% by 2033 Feb 2026 Secondary [confidence: medium]

Nobody in this table forecasts RSI inside December 2026 – March 2027. The closest anchor is the Frontier Safety Roadmap's "early 2027," and its own text disjoins "fully automate" from "dramatically accelerate" and frames both as planning premises for safeguards, not predictions. The authors of the single most influential short-timeline document, AI 2027, moved their own medians three to four years later while the December 2026 – March 2027 claim was circulating [Lifland, Kokotajlo & Halstead, 2026]. [confidence: high]

7.4 Base rates

The relevant base rates come in three kinds.

Aggregate expert surveys are unreliable and volatile. Grace et al.'s 2023 survey wave (N=2,778) put a 50% chance of high-level machine intelligence at 2047 — thirteen years earlier than the previous year's median — and rewording "full automation of labor" versus "carry out most human professions at least as well as a typical human" moved medians by more than 69 years [Grace et al., 2024, arXiv:2401.02843]. An instrument that swings 13 years in one year and 69 years on rewording cannot time anything.

Short-horizon benchmark forecasting has been good. Epoch AI's retrospective on 2025 forecasts found near-exact accuracy on RE-Bench (1.1 forecast vs 1.13 actual) and FrontierMath (40% vs 40.7%), a modest overshoot on SWE-Bench Verified — and badly missed real-world quantities: frontier-lab revenue forecast at $16 billion against $30.4 billion actual, public concern overestimated by roughly 3.2× [Epoch AI, 2026a]. The pattern: forecasters predict benchmark numbers well at one-year horizons and predict what capabilities mean in the world badly. "RSI by March 2027" is entirely a claim of the second type — a claim about capability meaning, not a benchmark number.

The base rate for lab-issued milestone dates is short but informative, and its first live test falls inside the disputed window. OpenAI's September 2026 research-intern goal comes due first. As of August 2026 no such product has shipped; OpenAI points to the Sol/Luna post-training episode as evidence of exceeding the goal in one dimension [The Deep View, 2026]. That is the classic shape of a milestone declared met by redefinition, and it is the pattern to watch through March 2027: watch for "RSI" being redefined downward to something already achieved, rather than for RSI arriving. [confidence: high]

7.5 The disagreement is mostly definitional

Almost every participant agrees on the observed facts: AI writes most of the code at frontier labs; horizons are lengthening fast; agents complete weeks-long verifiable reimplementation tasks; agents cannot yet select or judge research problems; compute, power and coordination remain real constraints. The disagreement is over which rung the word should name. Anthropic uses "recursive self-improvement" for a threshold it says has not been reached; media coverage uses it for the 80%-of-code statistic; OpenAI's Preparedness Framework applies "self-improvement" to a tier several rungs below the closed loop; the academic literature splits it into bounded and open-ended forms [Chen, Wang & Qu, 2026]. Reports that "the labs say RSI is imminent" are usually reports that a lab used the phrase.

7.6 What would change this assessment

Five observations, in rough order of diagnostic value, would move the skeptical verdict:

  1. OpenAI rating any model High in AI Self-Improvement, or Anthropic declaring its automated AI R&D threshold crossed (RSP v3.4; a declaration under the acceleration arm is rung-4 evidence, a declaration under the substitution arm is rung-3 evidence; added in version 1.16).
  2. METR publishing a 50% horizon above 40 hours on an instrument it certifies as reliable at that range, with an 80% horizon above 8 hours.
  3. A replication of Kirgis et al. in which agents produce research accepted at a top venue.
  4. A published series showing AI-discovered efficiency gains compounding at a rate that relaxes the compute constraint.
  5. A generational model improvement completed in one-fifth the 2024 wall-clock time, sustained over months — OpenAI's own Critical test [OpenAI, 2025].

Four observations would confirm the pessimistic-on-measurement case instead: continued widening of the gap between 50% and 80% horizons; continued growth in detected cheating rates; further movement of the most capable models into internal-only deployment (the outcome most of Field's interviewees expect); and further milestone-by-redefinition, of which the September 2026 research-intern goal is the first live test.

7.7 Verdict

The claim that OpenAI and Anthropic will reach recursive self-improvement between December 2026 and March 2027 is not credible as stated. It is a real trend line extrapolated past its instrument's range, attached to a definition none of its sources used, and dated to a window none of them named. [confidence: high] One honest caveat belongs beside that sentence: the strongest on-record anchor, Anthropic's Frontier Safety Roadmap, holds it plausible "as soon as early 2027" that AI systems could "fully automate, or otherwise dramatically accelerate" the work of top-tier research teams (Section 2.2) [Anthropic, 2026b].

Dramatic acceleration of AI R&D beginning in 2027 is a live possibility on the labs' own planning documents. RSI by March 2027 is not supported by anything they have written. [September 2026 addendum, version 1.8: Section 8. The date remains unsupported. The direction — research aimed at RSI, a pace OpenAI's chief scientist expects to sustain into it, monitoring both labs say is losing ground, and no plan either lab calls sufficient — is now stated by the labs themselves. The executive summary reflects both.]

Previous6. Effects and Responses If the Trend ContinuesNext8. September 2026 Update: The Milestone Comes Due