Recursive Self-Improvement at OpenAI and Anthropic · Section 2 of 9

2. Provenance: Where "December 2026 – March 2027" Comes From

Version 1.20, revised 28 September 2026

No primary, dated, on-record statement by OpenAI or Anthropic asserts that either company will reach recursive self-improvement between December 2026 and March 2027. Repeated independent tracing of the claim — across lab blogs, system cards, policy documents, and press coverage through August 2026 — reaches the same conclusion [confidence: high]. The window is a composite media artifact. Each component of the window traces to a real document; no document asserts the assembled claim. Four tributaries feed it.

2.1 The four tributaries

First tributary: the AI 2027 scenario. The specific March 2027 date traces most directly to "AI 2027," the narrative forecast published by the AI Futures Project in April 2025 (Kokotajlo, Alexander, Lifland, Larsen, Dean; ai-2027.com) [Kokotajlo et al., 2025]. The scenario places a "superhuman coder" at its fictional lab "OpenBrain" in March 2027, with an internal AI research assistant ("Agent-1") deployed around November–December 2026 and a superhuman AI researcher months later.

The match between the viral window and the scenario's internal dates is near-exact: December 2026 is when Agent-1 deploys, March 2027 is when the superhuman coder arrives. The authors present it as a modal scenario with wide uncertainty, not a lab prediction, and it is not a statement by OpenAI or Anthropic [confidence: high]. Secondhand claims that "labs say RSI by March 2027" re-date this scenario as if it were a lab plan — speculation-to-fact laundering.

The decisive update: the scenario's own authors have since moved their medians three to four years later. As of January 27, 2026, Kokotajlo's superhuman-coder median sits at approximately end-2029/early-2030 and his AGI median at December 2030; Lifland's moved to 2035. Kokotajlo: "Things seem to be going somewhat slower than the AI 2027 scenario" [Lifland, Kokotajlo & Halstead, 2026]. The authors no longer stand behind the date they supplied.

Second tributary: OpenAI's roadmap. On October 28, 2025 — the same day OpenAI completed its restructuring into a public benefit corporation — Sam Altman, chief scientist Jakub Pachocki, and co-founder Wojciech Zaremba stated internal goals of an "intern-level research assistant by September 2026" and a "legitimate AI researcher" by 2028, on a livestream [TechCrunch, 2025]. Neither date falls inside the December 2026 – March 2027 window, and neither is an RSI claim: "research intern" and "automated researcher" describe automation of research labor, not a closed loop in which an AI's contribution shortens the next AI's development cycle.

Secondary accounts render the 2028 date as specifically March 2028; that month appeared only in later re-reporting (TIME, MIT Technology Review) and was unconfirmed against the recording through version 1.5; OpenAI's own post of September 6, 2026 states the target as "an automated AI researcher by March of 2028" [OpenAI, 2026, Research acceleration; confirmed in version 1.6, see 8.10]. The window sits between OpenAI's two milestones — the September 2026 intern drifting later and the 2028 researcher drifting earlier meet in the middle. Where verifiable, OpenAI's own stated targets are later and weaker than the claim attributed to the company.

Third tributary: Amodei's capability statements. Dario Amodei has made a family of widely quoted claims that read, compressed, as an RSI date. At a Council on Foreign Relations event on March 10, 2025 (not Axios — press coverage conflated outlet with venue): "we'll be there in three to six months — where AI is writing 90% of the code. And then, in 12 months, we may be in a world where AI is writing essentially all of the code" [CFR recording].

In "Machines of Loving Grace" (October 11, 2024): "powerful AI," a "country of geniuses in a datacenter," possibly "as early as 2026" — an essay that also devotes extended analysis to physical-world bottlenecks and explicitly rejects an instant-singularity reading. At Davos in January 2025: systems "broadly better than almost all humans at almost all things" by "2026 or 2027" (a near-identical sentence appears in print in "On DeepSeek and Export Controls," January 2025).

The coding statements are rung-1/2 claims; the "powerful AI" statements are capability projections with no self-improvement mechanism; none uses the term RSI or names a closed loop. The compression of "90% of code" + "powerful AI 2026–27" into "Anthropic predicts RSI by early 2027" is the bait-and-switch in its most common form.

Fourth tributary: threshold laundering. Anthropic's Responsible Scaling Policy and OpenAI's Preparedness Framework define AI R&D capability thresholds — governance triggers that require safeguards if crossed — and these are routinely misread as forecasts. The threshold wordings, and the consequential RSP v3.0 restructuring, are given in Section 1. Crossing language of this kind describes a contingency the lab wants safeguards for, not a scheduled event. Reporting that converts "our policy covers models that could automate a junior researcher" into "the lab predicts automated researchers" commits the error — the clearest documented instance of definitional bait-and-switch in this discourse.

2.2 The strongest single anchor: the Frontier Safety Roadmap

One document comes close to the window. Anthropic's Frontier Safety Roadmap (July 10, 2026) states: "We believe it is plausible, as soon as early 2027, that our AI systems could fully automate, or otherwise dramatically accelerate, the work of large, top-tier teams of human researchers" [Anthropic, 2026b, https://anthropic.com/responsible-scaling-policy/roadmap]. This is the closest on-record lab statement to the claimed window, and it deserves to be stated fairly: a frontier lab, in its own name, in a dated document, describing early 2027 as a plausible date for research automation at team scale. If the claim under examination has a single strongest anchor, this is it.

Three qualifications keep it from carrying the claim. It is a safeguards-planning document — a statement about what Anthropic must be prepared for, in the same genre as the RSP, and using "plausible" rather than "expected." Its disjunction is doing heavy work: "fully automate, or otherwise dramatically accelerate" spans everything from rung 4 down to strong rung-2 assistance, and only the first disjunct approaches RSI. And it names no mechanism of recursive improvement — it is a claim about labor automation, not about a loop that shortens the next cycle. Read strictly, it supports "Anthropic plans for the possibility of dramatically accelerated research by early 2027," which is materially weaker than "Anthropic predicts RSI by March 2027." But anyone defending the claim window should cite this document.

2.3 Consolidated quote ledger

Speaker Verified wording Venue, date What it actually claims Verification
Amodei AI "writing 90% of the code" in 3–6 months; "essentially all" in 12 CFR event, Mar 10, 2025 Coding automation (rungs 1–2) CONFIRMED (venue corrected from Axios)
Amodei "country of geniuses in a datacenter"; powerful AI "as early as 2026" "Machines of Loving Grace," Oct 11, 2024 Capability projection, bottleneck-aware CONFIRMED
Amodei "broadly better than almost all humans at almost all things" by "2026 or 2027" Davos, Jan 2025 Capability projection Verified via transcripts
Amodei "models that were good at coding and good at AI research… to produce the next generation… and speed it up to create a loop" Davos, Jan 2026 RSI mechanism description, no date UNVERIFIED — single aggregated summary [confidence: low-medium]
Altman superintelligence possible "in a few thousand days" "The Intelligence Age," Sept 23, 2024 Capability projection (~decade+) CONFIRMED (venue corrected from "Three Observations")
Altman "We are now confident we know how to build AGI" "Reflections," Jan 6, 2025 Capability confidence, no date CONFIRMED
Altman "We are past the event horizon; the takeoff has started"; 2026 systems "that can figure out novel insights" "The Gentle Singularity," Jun 10, 2025 Gradual-compounding framing; novelty, not loop closure CONFIRMED
Altman, Pachocki, Zaremba "intern-level research assistant by September 2026"; "legitimate AI researcher" by 2028 OpenAI livestream, Oct 28, 2025 Internal product milestones (research labor) CONFIRMED via TechCrunch; "March 2028" confirmed primary by OpenAI, 6 September 2026 (8.10)
Pachocki "deep learning systems are less than a decade away from superintelligence" Same livestream Capability projection CONFIRMED via TechCrunch
Clark "recursive self-improvement has a 60% chance of happening by the end of 2028" X post, May 4, 2026 (https://x.com/jackclarkSF/status/2051312759594471886); "make a better version of yourself… completely autonomously" wording in Axios interview, May 7, 2026 Personal RSI probability estimate; ~30% by 2027 reported secondarily CONFIRMED (primary located)
Kaplan humanity decides "between 2027 and 2030" whether to let AI train itself Guardian, Dec 2, 2025 Decision-window framing, not a forecast CONFIRMED; the circulating "as little as a year away" version has NO primary source (social-media paraphrase)
Anthropic Institute RSI "is not inevitable" but "could come sooner than most institutions are prepared for"; weeks-long tasks in 2027 "When AI builds itself," Jun 2026 Capability projection + preparedness framing CONFIRMED
Anthropic "plausible, as soon as early 2027, that our AI systems could fully automate, or otherwise dramatically accelerate, the work of large, top-tier teams of human researchers" Frontier Safety Roadmap, Jul 10, 2026 Safeguards-planning threshold; closest statement to the window CONFIRMED
Anthropic / OpenAI RSP AI R&D thresholds; Preparedness "Critical" self-improvement threshold RSP v2.x/v3.0; PF v2, Apr 15, 2025 Governance triggers, not forecasts CONFIRMED
Kokotajlo et al. superhuman coder March 2027 at fictional "OpenBrain" "AI 2027," Apr 2025, ai-2027.com Scenario, not lab prediction; authors' medians now ~2030+ CONFIRMED

2.4 Incentive context

Every major lab statement above was made in proximity to fundraising or policy lobbying. Anthropic closed a $65 billion Series H at a $965 billion post-money valuation on May 28, 2026 and confidentially filed a draft S-1 on June 1, 2026 — days before the Anthropic Institute published "When AI builds itself" [Fortune, 2026a; TechCrunch, 2026; VERIFIED]. OpenAI reportedly closed a $120 billion round at $850 billion post-money in March 2026, with both companies reported targeting Q4 2026 listings [secondary; confidence: low-medium]. [Superseded, version 1.20: Altman ruled out a 2026 listing on September 12 and Anthropic's target moved to November; see 8.27.] The earlier statements cluster the same way: Amodei's CFR remarks weeks after Anthropic's March 2025 $3.5 billion raise; Altman's essays amid OpenAI's restructuring negotiations. Critics said so directly: Mark Riedl of Georgia Tech observed that "the big AI companies are all jumping on the 'recursive self-improvement' hype train" [Scientific American, 2026].

The discounting must be applied symmetrically. Acceleration claims serve capability leadership and valuation; doom-adjacent and preparedness claims serve safety positioning and regulatory strategy. Both directions are commercially loaded in this period, and incentive proximity is a reason for caution about every statement in the ledger, not a license to discard the ones that cut against a preferred conclusion.

2.5 The arithmetic signature and the verdict

A structural observation: the window is not arbitrary. It is what three unrelated artifacts produce when compressed. AI 2027's internal dates put Agent-1 at November–December 2026 and the superhuman coder at March 2027; the median METR extrapolation from the early-2025 data crosses the ~8-hour "human workday" autonomy threshold in the same late-2026-to-early-2027 range; and OpenAI's two roadmap milestones bracket it, with Clark's reported ~30%-by-2027 figure available to be read as a point prediction. The window carries an arithmetic signature: it falls out of the source documents' own numbers. That explains why it feels corroborated, and why the corroboration is illusory — the artifacts share inputs (the same METR trend, the same lab statements) rather than independently confirming a date.

The verdict on provenance: "December 2026 – March 2027" is a composite artifact. It is the AI 2027 scenario's milestone dates, plus Amodei's 2026–2027 capability language, plus OpenAI's coding-automation and roadmap statements, plus safeguard thresholds misread as forecasts, assembled through a definitional bait-and-switch that slides from "AI writes most of our code" through "AI automates AI research" to "AI designs its successor without humans" as if these were one claim with one date. The Frontier Safety Roadmap's "plausible, as soon as early 2027" line is the only lab statement that genuinely touches the window, and it is a planning document about automation-or-acceleration, not a prediction of recursive self-improvement. No lab has named a date inside the window. The scenario that supplied the dates has been walked back by its own authors.

Previous1. What RSI Would Have to MeanNext3. The Measured Evidence