Recursive Self-Improvement at OpenAI and Anthropic · Section 5 of 9
5. Bottlenecks and Takeoff Models
5.1 The formal takeoff models
Four formal or semi-formal models frame the takeoff debate; a body of growth economics pushes back on all of them.
Davidson's compute-centric model. Tom Davidson's "What a Compute-Centric Framework Says About Takeoff Speeds" (Open Philanthropy, 2023, https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/) treats effective compute and the fraction of AI R&D tasks that AI can perform as the drivers of progress [Davidson, 2023]. Its central estimates suggest the transition from AI that can do ~20% of AI R&D tasks to AI that can do ~100% could take only a few years, and that a software-only acceleration is possible — but the result is sensitive to the returns-to-software-R&D parameter and to how strongly the remaining non-automated tasks bottleneck the rest [Davidson, 2023]. [confidence: medium — a model output, heavily parameter-dependent, not a measurement.] The model explicitly builds in diminishing returns and bottleneck tasks; it is the principal formal case that software-only takeoff is plausible, not a naive hard-takeoff argument.
Aschenbrenner's "Situational Awareness." Leopold Aschenbrenner (June 2024, https://situational-awareness.ai/) extrapolates orders of magnitude of effective compute and algorithmic progress in a straight line to AGI by ~2027, then projects hundreds of thousands to millions of automated-researcher copies compressing a decade of algorithmic progress into a year [Aschenbrenner, 2024]. The critiques: the "unhobbling" gains being extrapolated are hard to measure and may not persist; the argument assumes compute for experiments does not bind, which Epoch AI contests directly [Epoch AI, 2023–2025]; and Aschenbrenner founded an AI investment fund for which the essay doubles as a thesis, so the incentive-weighting rule applies.
AI 2027 as a takeoff model. The AI Futures Project's "AI 2027" (Kokotajlo et al., April 2025, https://ai-2027.com/) formalizes its narrative through a timelines forecast and a takeoff forecast, and is the proximate source of the March 2027 date. As a model, its distinctive move is fitting a superexponential curve to the METR time-horizon series. The empirical critique of that curve fit — titotal's analysis, the authors' response, and the authors' own later median revisions — is treated in Section 4 and is not repeated here; the point that belongs in this section is structural: the model's dramatic dates are driven by the functional form and parameter choices, not by any bottleneck analysis, and the model omits hardware R&D automation entirely, which Eli Lifland himself later named as a reason his own model's takeoff runs too slow [Lifland, Kokotajlo & Halstead, 2026].
Erdil & Besiroglu on explosive growth. Ege Erdil and Tamay Besiroglu, "Explosive Growth from AI Automation: A Review of the Arguments" (arXiv:2309.11690, https://arxiv.org/abs/2309.11690), define explosive growth as roughly 30% annual global GDP growth — an order of magnitude above historical rates — and conclude that it "seems plausible with AI capable of broadly substituting for human labor, but high confidence in this claim seems currently unwarranted" [Erdil & Besiroglu, 2023].
Their emphasis is on Baumol effects and O-ring logic: if some essential tasks cannot be automated or scaled, those tasks become relatively more important and cap aggregate growth. Erdil & Besiroglu thus sit closer to the counterweight than to the fast-timeline generators. Epoch AI's later work extends the same skepticism to the software-only intelligence explosion specifically: even with AI research automated, compute for experiments and the pace of empirical iteration bound how fast gains can be realized [Epoch AI, 2023–2025]. [confidence: high — a stated Epoch position.]
5.2 The growth-economics counterweight
The Baumol argument applied to ideas. Aghion, Jones & Jones, "Artificial Intelligence and Economic Growth" (NBER w23928, https://www.nber.org/papers/w23928), show that when AI automates production, Baumol's insight generates sufficient conditions for balanced rather than explosive growth even under near-complete automation; applied to a model where AI automates the production of ideas, the same mechanism can prevent explosive growth [Aghion, Jones & Jones, 2017/2019]. Growth ends up governed by whatever essential task resists automation, not by the automated majority. In semi-endogenous versions, even fully automated research yields exponential rather than hyperbolic growth unless the effective researcher population itself explodes [Jones; Aghion, Jones & Jones, 2019]. This makes the RSI question a special case of a precise economic question: can AI automate idea production with no residual bottleneck task? Every constraint in Section 5.3 is a candidate residual task.
Ideas getting harder to find. Bloom, Jones, Van Reenen & Webb (AER 110(4), 2020, https://doi.org/10.1257/aer.20180338) document that holding outcome growth constant has required exponentially rising research inputs — doubling chip density today requires more than 18 times as many researchers as in the early 1970s, with research productivity in semiconductors declining roughly 6.8% per year [Bloom et al., 2020]. This is the base rate that any recursion must outrun. The skeptical reading is the sharp one: an automated researcher inherits the same idea-production function humans face. Automating the existing workflow multiplies inputs into that function; an explosion requires changing its curvature, and nothing in the benchmark record shows AI systems doing that. RSI must outrun declining marginal returns to research, not merely automate research [confidence: high — Bloom et al. is a landmark empirical result; its application to RSI is an interpretive extension].
The conservative anchor. Daron Acemoglu, "The Simple Macroeconomics of AI" (NBER w32487, https://www.nber.org/papers/w32487), estimates AI's total-factor-productivity effect at roughly 0.53–0.66% cumulative over a decade (conservative case below 0.53%, about 0.064% per year), on the argument that AI will meaningfully automate about 5% of work tasks in that window [Acemoglu, 2024]. [confidence: high — this is the published estimate; critics such as Korinek & Trammell (NBER w31815) object that it excludes new tasks and find singularity-type growth possible under full automation.] The spread between Acemoglu's decade-scale fraction of a percent and Erdil & Besiroglu's 30%-per-year threshold spans the entire debate; the December 2026 – March 2027 RSI window requires the far tail of that spread to be right within months.
5.3 The constraint stack, with 2026 numbers
A software-only intelligence explosion requires compute, wall-clock training time, serial research time, power, capital, chips, coordination, and taste to all be non-binding at once.
Experiment compute, and the elasticity that is not identified. AI research advances by running experiments, and experiments consume GPU-time; if automated researchers generate ideas faster than they can be tested, compute binds the loop [Epoch AI, 2023–2025; Davidson, 2023]. The key empirical parameter is whether research compute and cognitive labor are substitutes or complements. Whitfill & Wu fit constant-elasticity-of-substitution production functions to data from OpenAI, DeepMind, Anthropic and DeepSeek over 2014–2024 and got a split result: a baseline specification in which the two are substitutable (permitting an explosion) and a frontier-experiments specification in which they are complementary (blocking one) [Whitfill & Wu, 2025, arXiv:2507.23181]. The honest summary is that available data do not identify the parameter on which the whole software-only-explosion question turns.
Parallelization technology. Phil Trammell's Epoch AI report of July 29, 2026 (https://epoch.ai/publications/parallelization-constraints-could-delay-a-technological-singularity) introduces a parameter absent from standard takeoff models: "Effective research inputs are limited by whichever is scarcer: the raw research inputs; or the 'parallelization technology' needed to divide, execute, coordinate, and integrate their work" [Trammell, 2026].
Ten thousand copies of a competent agent are not ten thousand researcher-years if the technology to decompose problems and reintegrate results does not exist. Trammell offers no timing prediction and frames it as an open empirical question. The instruction-drift and resource-awareness failures documented in the Kirgis et al. shadow evaluation (Section 3.3) are direct empirical evidence that this coordination technology does not yet exist [Kirgis et al., 2026]. This is a direct technical objection to Aschenbrenner's millions-of-virtual-researchers argument.
Serial wall-clock time. A generational model improvement requires a training run, and frontier training runs take months regardless of how good the algorithm is; frontier 2026 runs are estimated at 1e26–1e27 FLOP [secondary infrastructure analysis, 2026; confidence: low-medium]. No amount of cognitive labor compresses a three-month run below the physical duration of the run, absent algorithmic changes that must themselves be validated by runs. OpenAI's Critical Self-Improvement threshold — a generational improvement in one-fifth the 2024 wall-clock time, sustained for months — is defined in wall-clock terms for exactly this reason, and no model has been asserted to meet it [OpenAI Preparedness Framework; METR, 2026c].
Power. Total datacenter critical IT power demand is projected to roughly double from about 49 GW in 2023 to 96 GW by 2026, with about 90% of the growth AI-related; power to train the largest frontier models is growing more than 2x per year, on trend to multiple gigawatts by 2030 [Epoch AI data-center research, 2026; confidence: medium]. Power interconnection and datacenter construction run on multi-year lead times [SemiAnalysis, 2024–2025]. This bounds how fast software gains can be cashed out into larger training runs, on a timescale of years, not the months a December 2026 – March 2027 loop requires. [confidence: high — physical lead times are well documented.]
Financing — currently not binding. Capital, at least, is not the near-term constraint. Epoch AI's August 12, 2026 analysis documents Anthropic announcing $50 billion in American compute infrastructure via vendor-supported structures — approximately $35 billion of debt for TPU systems and $15.2 billion of loans for 1.43 GW of critical IT capacity, with Broadcom backstopping $30 billion, on a platform designed to support more than 20 GW of frontier-lab deployment through 2028 [Hutcheson, 2026]. The same source reports Anthropic revenue growing from $9 billion at end-2025 to more than $47 billion by May 2026 [Hutcheson, 2026; UNVERIFIED — an extraordinary growth rate, not independently confirmed]. Hyperscaler capex projections for 2026 cluster in the $600–800 billion range [Credit Sights via secondary, 2026; confidence: low-medium].
Chips. Chip performance per dollar has grown an average of 49% per year, doubling roughly every 1.7 years [Epoch AI, 2026c; confidence: medium — Epoch data insight, not independently reconfirmed]. That is fast by any historical standard and nowhere near fast enough to make compute non-binding on a four-month loop. Rao's framing: "If each frontier iteration is gated by chips, fabs, memory… the loop cannot compound at the speed of thought" [Rao, 2026].
Hardware R&D automation. The channel the AI 2027 model omits is real but does not escape the physical clock. AI-assisted chip design is production-grade and incremental: Google's AlphaChip reinforcement-learning floorplanning has been used for TPU layouts since the 2021 Nature paper, and Cadence Cerebrus and Synopsys DSO.ai ship RL-based placement and routing as standard features in 2026 flagship EDA tools [agentic-EDA survey, arXiv:2512.23189; LLM-assisted EDA framework, arXiv:2601.14098]. AI compresses some chip-design steps from months to less, but respins cost tens of millions of dollars and roughly six-month delays, and fab capacity and interconnect and power lead times keep hardware iteration on a physical clock. This partially supports Lifland's adjustment (the models omit a real channel) while supporting the bottleneck case on timing.
Research taste. This is the bottleneck every party now concedes, including the parties with the strongest incentive not to. Anthropic's own text: research taste and judgment — "choosing which problems matter, which results to trust, and when an approach is a dead end" — remain "an area of human comparative advantage, for now" [Anthropic Institute, 2026]. The RSI survey finds "research direction-setting" prevents complete loop closure [Chen, Wang & Qu, 2026].
Jack Clark, from inside Anthropic: "There's a certain absence of valuable, intuitive creativity in today's AI systems" [MIT Technology Review, 2026a] — a bearish signal against his own 60%-by-2028 forecast. Kirgis et al. measured the gap as five specific failure modes (Section 3.3).
Arvind Narayanan supplies the externality version: for superintelligence to cure cancer, "the hard part is clinical trials requiring thousands of people and 10–15 years" — the bottlenecks are outside the computer, and "I don't think AI recursive self-improvement is going to magically obviate those bottlenecks" [Narayanan, ICML 2026 keynote, via secondary; TIME, 2026]. If taste is a bottleneck task in the O-ring sense, the rate-limiting step of the loop stays in human hands even as coding agents improve [Erdil & Besiroglu, 2023; Aghion, Jones & Jones, 2019].
5.4 Where the bottleneck case is weakest
The bottleneck argument assumes the current research paradigm, and there is direct evidence of software relaxing the constraints it names. AlphaEvolve's fleet-scheduling heuristic recovered approximately 0.7% of Google's compute — roughly 14,000 servers — a case of AI-generated software directly loosening a compute constraint [Google DeepMind, 2025/2026]. Codex analyzed weeks of production traffic and wrote GPU-partitioning heuristics that raised token generation speeds by more than 20% ahead of the GPT-5.5 launch [OpenAI, 2026b; confidence: medium — vendor self-report, no independent replication].
If gains of this kind compound, "compute is binding" weakens over time, and the constraint stack becomes a moving target rather than a wall. The honest position is that the compounding rate of AI-discovered efficiency gains is unmeasured: the labs have not published the time series (efficiency gains attributed to AI-generated work, release over release) that would settle whether these are occasional harvests or a curve. That series is the single measurement that would most directly arbitrate between the bottleneck case and the takeoff models.