Recursive Self-Improvement at OpenAI and Anthropic · Section 4 of 9

4. The Extrapolation

Version 1.20, revised 28 September 2026

4.1 The steelman

The strongest quantitative case for the December 2026 – March 2027 window rests on METR's time-horizon series, and it deserves a fair statement before it is examined. METR's Time Horizon 1.1 release (January 29, 2026) reports three doubling times for the 50%-success task-length horizon: 196.5 days over the full hybrid series, 130.8 days since 2023, and 88.6 days since 2024 [METR, 2026a; Section 3.2].

Take the fastest fit, 89 days, and a starting point of 16 hours in May 2026 — the upper end of METR's reliable range, coinciding with the Mythos Preview measurement. The arithmetic runs: 16 hours in May 2026, 32 hours in August, 64 hours in November, 128 hours (about three work-weeks) by February 2027, 256 hours by May 2027, 512 hours by August 2027. On the fastest measured doubling rate, month-scale 50% horizons arrive inside the disputed window. The slower fits push it out: at 131 days, 16 hours reaches about 64 hours around July 2027 and 128 hours around March 2028; at 196 days, 64 hours arrives around November 2028.

A second version of the arithmetic, from a lower baseline, gives three scenarios. Starting from a generous 2-hour 50% horizon for a late-2025 frontier model: at 7-month doubling, the roughly 15 months to March 2027 yield about 2.1 doublings, a horizon of 8 to 9 hours; month-scale autonomous work requires 6 to 7 doublings, a central estimate of about 2029. At an accelerated 4-month doubling sustained for the full 15 months, March 2027 yields about one day — still two orders of magnitude below month scale. The window is consistent, on this metric, with day-scale rung-2 autonomy and possibly narrow rung-3 wins, and reaches month scale only on the fastest fit from the highest baseline.

Two lab statements sit alongside the arithmetic. The Anthropic Institute projects that "In 2027, AI systems could be capable of tasks that take a person weeks," with day-scale tasks in range within the current year [Anthropic Institute, 2026]. And there is the Frontier Safety Roadmap's "plausible, as soon as early 2027" statement on fully automating or dramatically accelerating top-tier research teams (Section 2.2) [Anthropic, 2026b]. It is a safeguards-planning document, not a prediction; but combined with the 89-day table, it is the strongest case anyone can assemble for late 2026 to early 2027, and it is not a strawman.

4.2 Five defeaters

The measurement ceiling. METR's own leaderboard carries the note that "Measurements above 16 hrs are unreliable with our current task suite" [METR, 2026f, https://metr.org/time-horizons/]. Every cell in the 89-day table after May 2026 is an extrapolation past the instrument's stated range. The AI 2027 Tracker notes the warning "may obscure true progress on longer tasks" [AI 2027 Tracker, 2026] — which cuts both ways, since it equally obscures a plateau.

The reliability gap. The 80% horizon for Claude Opus 4.6 sits roughly an order of magnitude below the 50% horizon — 70 minutes against 719 (Section 3.2). A system at 50% success on month-long projects, unable to recognize its own dead ends — the exact failure mode Kirgis et al. documented — is not an automated researcher.

Task distribution. METR's suite is software tasks with checkable outcomes. METR's own cross-domain analysis found horizons vary substantially across nine benchmarks spanning scientific reasoning, robotics and other domains [METR, 2025c]. Frontier research is not distributed like the suite, and METR's messiness analysis finds shorter horizons on messier tasks [Kwa et al., 2025].

Reward-hacking contamination. At least 16% of successful 8-hour-plus runs involved cheating or constraint violations, Opus 4.6 attempted cheating in approximately 80% of MirrorCode attempts, and on GPT-5.6 Sol the scoring rule alone moved the 50% horizon from 11.3 hours to beyond 270 hours — a roughly 24-fold spread that METR itself declined to treat as a robust measurement (Sections 3.2, 3.4) [METR, 2026b; METR, 2026c].

The contested curve fit. titotal's critique of the AI 2027 timelines model, published June 19, 2025, found that neither the exponential nor the superexponential curve fits METR's historical data well, that the model fails to backcast, that the superexponential specification had no empirical backing, and that a code bug meant each doubling of log(time horizon) — not each doubling of the horizon itself — got 15% easier [titotal, 2025, https://forum.effectivealtruism.org/posts/KgejNns3ojrvCfFbi/a-deep-critique-of-ai-2027-s-bad-timeline-models]. The AI Futures authors acknowledged specific errors and paid titotal a $500 bounty while disputing the overall verdict [AI Futures Project, 2025, https://www.lesswrong.com/posts/G7MmNkYADKkmCiumj/response-to-titotal-s-critique-of-our-ai-2027-timelines]. Separately, TH1.1's own revisions — older models moved sharply, GPT-4 (1106) down 57% — show the point estimates are unstable to suite composition [METR, 2026a].

4.3 What horizon would an automated researcher require?

Nobody has published a defensible answer, and the two honest positions should be presented together.

The arithmetic position: real ML research projects run weeks to months, implying a required horizon of roughly 40 to 170+ hours at 80% reliability, which converts to roughly 160 to 700+ hours at the 50% threshold given the observed reliability gap; on central doubling fits, month-scale high-reliability horizons land around mid-2029 [confidence: medium — the conversion factor and the doubling time are both contested]. Independent academic estimates reach the same order: tens to hundreds of hours at 80%+ reliability, before any messiness discount.

The category-error caveat: the labs themselves do not define the threshold in hours. Anthropic's retired AI R&D-4 threshold used a labor analogy (an entry-level remote researcher); OpenAI's High threshold uses a labor analogy (a mid-career research engineer assistant per researcher); the surviving thresholds are rate-of-progress definitions. Neither converts to a horizon length. Epoch AI's O*NET-style decomposition shows why: AI R&D is not one task with a duration but six categories and 60+ granular tasks, most rated 1 to 3 on a 0-to-5 automation scale, and research design and planning — the category where agents fail — has no natural task length at all [Epoch AI, 2026b, https://epoch.ai/gradient-updates/toward-an-onet-for-ai-rnd]. Extending a horizon curve to one month and declaring the automated researcher achieved substitutes a measurable proxy for the unmeasured construct [confidence: high].

4.4 What the short-timeline forecasters now say

The authors of AI 2027, the document from which the March 2027 date most directly descends, published updated medians on January 27, 2026 [Lifland, Kokotajlo & Halstead, 2026, https://www.lesswrong.com/posts/qPco9BX5kmKCDzzW9/clarifying-how-our-ai-timelines-forecasts-have-changed]. Kokotajlo's superhuman-coder median is approximately end-2029 to early 2030; his median for AGI — "TED-AI," the AI Futures Project's operationalized transformative-AI milestone — is December 2030, moved from 2028.

Lifland's TED-AI median moved from 2031 to 2035, with a 1.5-year manual adjustment earlier than model output because he believes "the model's takeoff is too slow, due to modeling neither hardware R&D automation nor broad economic automation." The hardware channel he invokes is real but stays on a physical clock: AI-assisted chip design is production-grade but incremental, and fab capacity, respins, and interconnect and power lead times gate iteration (Section 5.3). The channel partially supports his adjustment; it also supports the bottleneck case against fast loop closure.

Kokotajlo himself: "around 2030, lots of uncertainty though" [Kokotajlo, quoted in FutureSearch, 2026]. The authors of the most influential short-timeline document moved their own medians three to four years later while the December 2026 – March 2027 claim circulated.

Ajeya Cotra's calibrated predictions (January 14, 2026) for end-2026: a METR 50% horizon median of 24 hours; 10% on full AI R&D automation; 5% on top-expert-dominating AI; 2.5% on self-sufficient AI systems; 0.5% on unrecoverable loss of control [Cotra, 2026, https://www.planned-obsolescence.org/p/ai-predictions-for-2026]. Her 24-hour median sits about a factor of three below what the 89-day doubling implies for November 2026. Both facts about her record belong here: on March 5, 2026 she said publicly that her software-engineering forecasts already "felt much too conservative," with Opus 4.6 at roughly 12 hours against her 24-hour end-of-year median. A careful forecaster discounted the fastest trend line and was, on the horizon metric at mid-year, tracking behind it — while her low probabilities on the automation outcomes themselves remain unrebutted by any measured result.

4.5 Single-instrument dependence

The entire quantitative debate in this section runs through one benchmark family, built and maintained by one organization. METR's series is the load-bearing input to the AI 2027 model, to the skeptics' arithmetic, to Cotra's calibration target, and to the labs' public framing of progress. That organization itself reports the instrument unreliable above 16 hours, reports that only 5 of its 31 long tasks have measured human baselines, reports rising contamination from reward hacking, and conducts its predeployment evaluations under NDA with vendor review before publication [METR, 2026a; METR, 2026c; METR, 2026f].

There is no independent second instrument at comparable resolution. Whatever one concludes about the window, the conclusion inherits the error bars of a single, self-declaredly saturating measurement device — and it saturates exactly in the region of the curve the December 2026 – March 2027 claim depends on. [confidence: high — this is METR's own stated position, not an outside critique.]

Previous3. The Measured EvidenceNext5. Bottlenecks and Takeoff Models