Recursive Self-Improvement at OpenAI and Anthropic · Section 1 of 9

1. What RSI Would Have to Mean

Version 1.20, revised 28 September 2026

The proposition "OpenAI and Anthropic will reach recursive self-improvement between December 2026 and March 2027" cannot be evaluated until "recursive self-improvement" is pinned down. The term is used across at least five distinct meanings in lab communications and the literature, and the credibility of the timeline claim varies by orders of magnitude across them. The most common rhetorical move in this discourse is a definitional bait-and-switch: sliding between "AI writes most of our code," "AI automates AI research," and "AI designs its successor without humans" as though they were a single claim.

The concept itself originates with I.J. Good: "Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion'" [Good, 1965]. Good's formulation contains the essential recursion: the machine improves the process that produces machines, so gains compound.

The five-rung ladder

Every timeline claim in this report is assessed rung by rung. The rungs are not interchangeable, and evidence for one is not evidence for the next.

Rung 1 — AI-assisted coding at labs. Humans direct; AI writes code. Verified and saturated as of 2026.

Rung 2 — Autonomous multi-hour engineering. An agent completes a bounded, well-scoped software or ML engineering task end-to-end at human-parity reliability, with limited intervention. Partially verified; the measurement instruments are contested (Section 3).

Rung 3 — Autonomous generation and validation of novel research ideas yielding real algorithmic gains. The agent chooses the idea, not just the implementation, and produces gains a competent human researcher would recognize as real contributions. Weakly and narrowly verified.

Rung 4 — A closed loop. AI-produced gains measurably shorten the cycle time of the next round of AI development — R&D output becomes an input to R&D speed. This is the minimal operational definition of "recursive" self-improvement: the loop closes. Not demonstrated at frontier scale.

Rung 5 — Sustained superexponential improvement with humans out of the loop. The classic intelligence explosion of Good, Yudkowsky, and Bostrom [Bostrom, 2014]. Not demonstrated; no lab claims it.

Traced to primary sources, the December 2026 – March 2027 claims almost always concern rungs 1–3. Genuine RSI as most people intuitively understand the term — and as it drives the safety and geopolitical stakes — requires rung 4 at minimum, and rung 5 for hard-takeoff scenarios. A recurring failure mode in the public discourse, documented in Section 2, is sliding from rung 1 to rung 4 within a single sentence.

Soft versus hard takeoff

A second distinction cuts across the ladder. Compounding ("soft") RSI: improvement is real and self-reinforcing but bounded by bottlenecks and diminishing returns — each round of AI-assisted research is gated by compute, experiment wall-clock, and declining marginal returns to ideas, so the trajectory is fast exponential or mildly superexponential growth rather than a singularity. Epoch AI's modeling and Erdil & Besiroglu's work on explosive growth fall broadly into this camp: they take accelerated growth seriously while emphasizing complementary bottlenecks [Erdil & Besiroglu, 2023, arXiv:2309.11690; Epoch AI, 2023–2025]. Hanson and Davidson's compute-centric takeoff model treat RSI as continuous acceleration of effective-compute growth [Davidson, 2021–2023, Open Philanthropy].

Hard-takeoff RSI: a discontinuous, fast bootstrap to vastly superhuman capability, plausibly within days to months, with humans unable to intervene — the scenario emphasized by Yudkowsky, in some readings of Bostrom, and in the faster branches of AI 2027 [Kokotajlo et al., 2025, https://ai-2027.com].

The distinction matters for policy, not just taxonomy. Compounding RSI implies a fast but potentially governable transition: pre-deployment evaluation, capability thresholds, and human oversight retain traction because each cycle still runs through observable training runs and deployments. Hard takeoff implies those instruments could fail catastrophically and quickly, because the transition outruns the evaluation cadence. Dean Ball splits the outcome space along the same seam: a "normal exponential" case in which RSI means faster progress within the existing paradigm, largely invisible to the public, versus a case in which the nature of the technology changes discontinuously [Ball, 2026]. Most public argument conflates the two.

The labs' own operational definitions

Both labs have written down thresholds, which is more useful than any interview quote because thresholds carry governance consequences. Both are safeguard triggers, not predictions — reading them as forecasts is a common provenance error in this debate ("threshold laundering," Section 2).

OpenAI Preparedness Framework v2 (April 15, 2025) tracks AI Self-Improvement as one of three frontier categories, with two thresholds [OpenAI, 2025, https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf; wording verified]:

  • High: the model is "capable enough that it is equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant, relative to those researchers' 2024 baseline."
  • Critical: either (leading indicator) a superhuman research-scientist agent, or (lagging indicator) causing "a generational model improvement (e.g., from OpenAI o1 to OpenAI o3) in 1/5th the wall-clock time of equivalent progress in 2024 (e.g., sped up to just 4 weeks) sustainably for several months."

The Critical definition is an unusually precise public operationalization of RSI, because it is measured in wall-clock compression of generational improvement sustained over months. It sits at rungs 4–5. High sits at roughly rungs 2–3.

Anthropic's Responsible Scaling Policy was restructured in v3.0, effective February 24, 2026 [Anthropic, 2026a, https://www.anthropic.com/responsible-scaling-policy]. Under v2.x there were two AI R&D thresholds: AI R&D-4, "the ability to fully automate the work of an entry-level, remote-only Researcher at Anthropic," and AI R&D-5, "the ability to cause dramatic acceleration in the rate of effective scaling."

RSP v3.0 collapsed these into a single automated AI R&D threshold. Its wording was: "Our working operationalization is to trigger this risk threshold at the point where we determine that a model could compress two years of 2018 – 2024 AI progress into a single year" [Anthropic, 2026a, https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0; wording verified against the PDF]. This effectively adopted the former AI R&D-5 definition, and the v2.x commitment to develop an "affirmative case" at the AI R&D-4 level was removed [GovAI, 2026, https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections].

That wording lasted five weeks. RSP v3.1, effective April 2, 2026, deleted the sentence and put a two-part test in its place. The threshold is met "if we determine that either (1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is “dramatic acceleration” of the pace of AI progress for reasons that likely relate to the automation of AI R&D" [Anthropic, 2026, RSP v3.1 redline, https://cdn.sanity.io/files/4zrzovbb/website/64cb0ac5eb0f8030187131f490827323e3d53308.pdf]. Anthropic's web changelog described the v3.1 edits as ones "neither of which significantly change the substance of the policy" and mentioned only the clarification of "doubling"; the changelog inside the PDF says the change "involves some substantive changes to make it better reflect our underlying intent" [Anthropic, 2026, https://www.anthropic.com/responsible-scaling-policy; Anthropic, 2026, RSP v3.4, Changelog]. Neither names the substitution arm.

RSP v3.4, effective July 8, 2026, is the current version, and it tightened the second arm. Scenario (2) has occurred where "(a) we observe or expect double the rate of progress in AI aggregate capabilities compared to both the rate we’d expect and the fastest rate of extended progress we’ve observed in the absence of significant AI contributions to AI R&D and (b) it is plausible that this doubling is substantially attributable to the automation of research and/or engineering (as opposed to other factors, such as increased headcount, compute, or general productivity), such that continuation of the trend in AI progress seems likely to lead to even greater acceleration" [Anthropic, 2026, RSP v3.4, Section 1, pp. 9–10, https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf]. A footnote defines the doubling as "as much progress in one year as one would see in two years at baseline" and adds that this "is not the same idea as 'doubling researchers’ productivity.'" A second footnote defines "extended" as "Over at least three model generations." The words "both" and "the fastest rate of extended progress we’ve observed" are the v3.4 insertions. The changelog gives the reason: if overall progress were "constant or slowing" while Anthropic estimated it was much faster than it would have been without AI tools, the threshold would not be crossed. The same entry says: "This threshold is intended to capture the onset of dramatic recursive self-improvement, and has proven difficult to operationalize" [Anthropic, 2026, RSP v3.4, Changelog, July 8, 2026].

The governance change of February stands and has been underdiscussed: the threshold that would have fired at approximately rung 3 was retired and has not been restored. Anthropic separately committed to meeting the AI R&D-4 standard for all future models exceeding Claude Opus 4.5's capabilities rather than making case-by-case capability judgments [confidence: medium — secondary summary of the Opus 4.6 system card].

On this report's ladder the two arms sit in different places. The second arm is rung 4 in the lab's own words. It requires a measured or expected doubling of aggregate capability progress, attributed to automation and not to headcount or compute, of a kind that "seems likely to lead to even greater acceleration." That last clause is the closed loop: R&D output becoming an input to R&D speed. Anthropic measures it as a rate of capability progress and this report's rung 4 is stated as cycle time, but the two describe the same event, and Anthropic's Risk Report calls the doubling "a potential early warning" of "super-exponential progress," which is rung 5 [Anthropic, 2026, Risk Report, August 2026, Section 3.5]. The v3.4 edit moved this arm closer to rung 4, because a doubling that exists only against a counterfactual no longer counts.

The first arm is on no rung exactly. It is a statement about capability and cost, and it makes no reference to the pace of progress. Substituting for every Research Scientist, senior staff included, requires rung 3 across the whole range of research work, research taste included; Anthropic's own assessment is that its models "do not yet substitute for our Research Scientists and Research Engineers, especially relatively senior ones" [Anthropic, 2026, Risk Report, August 2026, Section 3.4]. A lab in that position would have the capability rung 4 needs and would probably reach rung 4 soon after, but the arm can be declared met before any acceleration is measured. This report therefore places arm (1) at the top of rung 3, as a sufficient capability for rung 4 and not a demonstration of it. A declaration under arm (1) alone would not satisfy the rung-4 bar; a declaration under arm (2) would.

Note the asymmetry between the two labs. OpenAI's High is a researcher-assistant threshold. Anthropic's threshold is met either by a rate of progress or by replacement of the whole research staff, and both arms are far above an assistant for each researcher. A model could plausibly sit above OpenAI's High and below both arms of Anthropic's threshold simultaneously. Cross-lab comparisons of "how close are we" are therefore not like-for-like, and any claim that names both labs and one date should be read with that in mind.

[Correction, version 1.16: versions 1.0 to 1.15 described Anthropic's threshold in its RSP v3.0 wording only, quoted through a secondary source. RSP v3.1 (April 2, 2026) replaced that wording with the two-part test quoted above, and RSP v3.4 (July 8, 2026) tightened its second arm. Version 1.15 (8.18) reported the change from Anthropic's August Risk Report without opening the policy. This revision read RSP v3.0, v3.1, v3.2, v3.3 and v3.4 and the v3.1 and v3.4 redlines from the primary PDFs. Open question Q13 is closed.]

The academic taxonomy

A July 2026 survey distinguishes "bounded self-refinement" — convergent, measurable, already in industrial use — from "open-ended recursive self-improvement," and organizes systems along two axes: what improves (deployment behavior, training policy, evaluators, or the research process itself) and degree of loop closure, from human-supervised to fully autonomous [Chen, Wang & Qu, 2026].

Its central empirical finding is that demonstrated self-improvement strength tracks a verification hierarchy: strongest where formal verifiers exist, weakest where the system assesses itself, with failure modes including self-confirming loops and model collapse. Its bottleneck conclusion: "research direction-setting" remains a domain requiring human involvement across frontier labs, preventing complete loop closure [Chen, Wang & Qu, 2026].

This verification-hierarchy finding recurs throughout the evidence sections — every strong 2026 self-improvement result sits at the formal-verifier end of the hierarchy, and coding is the near-term channel precisely because it is the domain where verification is cheapest.

The standard applied in the rest of this report follows directly: the December 2026 – March 2027 claim is held to the rung-4/5 bar under the labs' own written definitions — OpenAI's Critical threshold and the acceleration arm of Anthropic's automated AI R&D threshold — with rung-1 through rung-3 evidence credited as exactly what it is and no more.

PreviousOverviewNext2. Provenance: Where "December 2026 – March 2027" Comes From