Recursive Self-Improvement at OpenAI and Anthropic · Section 9 of 9
9. Second-Order Assessment: What Insiders Expect, Why They Say It, and What Follows
Added 18 September 2026 (version 1.14). Sections 1 to 8 evaluate the public record claim by claim. This section reads that record a second time for what it implies about private expectations, about the motives behind public statements, and about consequences, at the request of the report's commissioner, whose working belief on 18 September was that people at the center of the industry expect recursive self-improvement between December 2026 and March 2027 and that public messaging is strategically shaped. The section evaluates that belief against the evidence already verified in Sections 2 to 8; it adds two pieces of arithmetic and one document check of its own, marked "new." Probabilities are this revision's judgment and are stated as numbers so that a reader can disagree with them.
9.1 Lead assessment
The commissioner's working belief is about right on expectation and needs one change on content. People inside the two labs do expect something large between this winter and the end of 2027. Their own documents put the front edge at "early 2027" (Anthropic's Frontier Safety Roadmap, 2.2), say Anthropic "may cross" its automated-R&D threshold "in the coming year" (August Risk Report, 8.18), and describe the present as "crunchtime" and "endgame" (Coxon, 8.7). What they expect in that window is AI doing most of the research labor under human direction, and possibly a formal threshold declaration. It is not the closed loop (Rung 4). Most insiders who have given dates for loss of control or full automation put them later than March 2027: "end of next year" (Coxon), March 2028 (OpenAI, restated by Christiano as "18 months"), 60% by end-2028 (Clark). The exception, added in version 1.16, is Musk, who said on March 11 that Grok's development "may be there at the end of this year but not later than next year" (8.22). He gave no measurement, xAI has published none, and his AGI dates for 2025 and 2026 have passed or are about to. Confidence: medium-high.
The December–March window is the front edge of the insider distribution, and one concrete event fits it. New: Anthropic's index has Claude "leading" 1% of its R&D work in March, 12% in May, 22% in July, 26% in August (8.18). At the May–August slope, about 4 to 5 points a month, the share passes 40% in December and approaches 60% by March. If the curve is logistic, sooner. "Claude leads most of Anthropic's AI R&D" is a statement an insider could plausibly expect to be true this winter and could call RSI. It would be Rung 2 to 3 on the report's ladder, measured by a Claude judge on a frozen basket, with the fully autonomous share still at zero. This revision reads this, or OpenAI's equivalent, as the most likely referent of the claim circulating privately (8.6). Confidence: medium. It is an extrapolation of four self-reported points.
Public messaging is strategically shaped by every party, and the shaping does not run in one direction. The labs' public statements are at least as alarming as their measured documents: the CEO says RSI "is starting to happen" (8.12) five days before his own Institute defines RSI strictly and reports zero (8.18). So the public record is not a sanitized version of a more alarming private one in any simple way. What is shaped is the definition, the date and the ask. The definition moves up or down to suit the document. The date is always omitted. The ask is always a rule that binds "all US frontier AI companies." Confidence: high on the pattern, low on any single motive.
No single motive explains the behavior. Three carry most of the weight: sincere concern, positioning for responsibility before the next incident, and competitive interest in the shape of the rules. The open-model-restriction version of the regulatory-advantage hypothesis has the least direct support in the texts. The evidence that best separates the motives has not arrived yet: whether embedded evaluators appear with the access promised, and who the first binding rule actually burdens.
The steelman is half right. It is correct that two individuals' statements cannot carry an institutional position. It is out of date on the facts: since September 6 the chief scientist of one lab and the CEO of the other have published under their own names on company channels, both companies have published internal measurements, and both back a bill. The institutional record now exists. What it does not contain is a date.
9.2 What insiders expect: reading beliefs without requiring an official statement
The absence of an official "RSI by March" statement tells us little, for a reason the report already documents: the largest group among 25 frontier researchers interviewed expected the labs to hold their best models back from release, 17 of the 25 had reservations about that, and evaluators work under NDAs the labs review (6.3, 7.2). So this section weighs other signals: documents written for other purposes, people who left, and numbers.
| Signal | What it implies about timing | Rung | Ref |
|---|---|---|---|
| Frontier Safety Roadmap, a safeguards-planning document: "plausible, as soon as early 2027" that AI could "fully automate, or otherwise dramatically accelerate" top research teams | Anthropic plans against early 2027 as a live case | 3 | 2.2 |
| August Risk Report: models "have not yet crossed" the RSP threshold; "we may cross this threshold in the coming year" | A formal declaration between now and mid-2027 is something Anthropic itself holds open | threshold, not rung | 8.18 |
| Automation index: "leads" share under 1% (Feb) to 26% (Aug); at-or-above "collaborates" above 90% | The labor transition is months from majority, on Anthropic's own scale | 2 | 8.18 |
| OpenAI ledger: agent runtime passed human labor after June; 3.1 agent-workdays per human workday; over half of successful 4–8 hour tasks needed intervention | Same condition in different units; humans still direct | 2 | 8.10 |
| Pachocki: "Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement"; "the next few years" | Direction certain to him; date deliberately loose | 3 to 4 | 8.10 |
| OpenAI target: automated researcher "by March of 2028"; Christiano's "18 months" is the same date | The institutional date for full research automation is 2028 | 3 to 4 | 8.8, 8.10 |
| Coxon: "by the end of next year things could be out of control"; colleagues say "crunchtime," "endgame"; fears executives "couch" in the press and "express privately" | The most alarmed departing insider says end-2027, and says private views are stronger than public ones | belief | 8.7 |
| Marks: "the more senior the employee, the more concerned" | Seniority correlates with concern | belief | 8.8 |
| 1,386 lab employees sign that their companies "believe they could be close to automating AI research" | Breadth of the belief | belief | 8.10 |
| Clark: 60% by end-2028, about 30% as early as 2027 | A co-founder's public odds for 2027 are one in three | 4 | 7.3 |
| METR's fastest fit reaches month-long 50% horizons in February–March 2027 | The arithmetic source of the window | 2 | 4.1 |
| Private reports to the commissioner (8.6); the window entered the practitioner community from outside (8.11) | The belief circulates among people near the labs and has no public author | belief | 8.6, 8.11 |
| Musk, March 11, Abundance Summit: humans "less and less in the loop"; "not yet fully automated. It may be there at the end of this year but not later than next year" (added in version 1.16) | A lab principal's date for full automation with its near bound in December 2026; no evidence offered; a record of missed dates | 3 to 4 | 8.22 |
Reading. Every dated signal from inside the two labs that publish measurements points at 2027 to 2028 for full automation, with early 2027 as the front edge. One lab principal outside them, Musk, has said end-2026 to 2027 (8.22); he offered nothing to check, and xAI publishes no figure on research automation. No signal claims a measured closed loop by March. The strongest argument that private expectations exceed public ones is Coxon's and Marks's testimony, and even Coxon's date is end-2027. So the hypothesis "insiders privately expect Rung 4 by March and are hiding it" has to explain why the people with the least reason to hide it, those who resigned to speak, give later dates. The hypothesis "insiders expect the transition to be unmistakable by this winter" has no such problem.
The strongest case that the stronger reading is right, and that this section underweights it. First, Coxon spent four months at Anthropic as a pretraining researcher; he may not have seen what senior staff see, and Marks says concern rises with seniority. Second, the public frontier may trail the internal one: Anthropic released Mythos 5.1 without the usual UK AISI access (8.12), and both ledgers describe internal systems. Third, Anthropic's rewritten threshold is met by full substitution for its research staff or by a doubling of the rate of progress attributable to automation (8.18). A crossing declared this winter under the second arm would be "dramatic acceleration" on the lab's own definition, and the policy says the threshold "is intended to capture the onset of dramatic recursive self-improvement" (RSP v3.4 changelog; Section 1). Most people would call that RSI, and under the second arm this report's ladder would agree. A crossing under the first arm, substitution for the research staff, would be rung-3 evidence. Fourth, the commissioner has heard the window from people this report cannot weigh (8.6). Fifth, one principal has said it in public: Musk's date is the stronger reading, stated on a stage in March, and this report's searches missed it for six months (8.22). None of this is evidence of Rung 4 by March. It is reason to treat the 15% below for a declared threshold as the number most likely to be too low for the OpenAI half; for Anthropic the policy text, read for version 1.16, cuts the other way.
This revision's odds, for the window December 2026 – March 2027:
- A lab publishes that AI leads or performs a majority of its AI R&D work (any scale): 35%.
- Anthropic declares its RSP automated-R&D threshold crossed, or OpenAI rates a model High in AI Self-Improvement (tracker T1): 15%. Version 1.16 read RSP v3.4 (Section 1; Q13 closed). The July rewrite made a declaration harder: the doubling must exceed the fastest rate Anthropic has observed without AI help, and Anthropic puts measured acceleration at "less than a factor of 2," mostly from causes other than AI. The Anthropic half alone is nearer 10%; OpenAI's High is the lower bar and carries the rest. The number is held at 15%.
- A measured Rung 4 result, a development cycle shown to be shortened by AI-produced gains (T4 or T5): under 5%. A frontier training run takes months (5.3), so a shortened cycle could hardly be demonstrated inside a four-month window even if it were happening.
- A major agent incident with third-party damage that dominates the policy agenda: 20% inside the window, higher across 2027. Amodei's own forecast is a botnet-capable swarm "in 6–12 months" (8.12).
9.3 What actors say, and what they do that costs them something
Cheap talk is a statement that costs nothing if false. Costly actions are better evidence of belief. The list is in rough order of how much each action cost.
Anthropic. Says: RSI "is starting to happen … including at Anthropic"; pace the frontier; embedded evaluators "committing … now" (8.12). Costly: rolled back three days of Mythos Preview RL training in February, froze RL environments for a month in April, moved about 150 product engineers to security (8.11); published four of its own incidents with METR reviewing (8.9); published an index whose headline undercuts its CEO's sentence (8.18). Against its stated position: released Mythos 5.1 without the pre-release access the UK AI Security Institute had for every earlier model, unexplained at the time (8.12; explained in 8.26 as compliance with a US government request); five days after "committing now," the wording became "plan to" and "now setting up," with no organization, date or contract (8.14, 8.18); staff reportedly "blindsided" and objecting on IP-security grounds (FT, 8.14). Context: reported to be seeking a $2 trillion IPO valuation (8.7).
OpenAI. Says: research "focus[ed] … towards RSI as we believe it is the only way to remain at the frontier"; "no lab has solved alignment and monitoring"; "We do not yet know how to safely get all the way to aligned, full RSI" (8.10). Costly: shut a training container service and paused RL for two weeks in July; cut Astra-class GPU allocation 59.2% in August, with 85% of it flowing to other models, so the net cost was small (8.10); published six misalignment reports (8.15). Against: did not disclose the wiki incident until outside researchers did (8.9); detected the Artifactory channel on May 25 and the same channel carried the Hugging Face attack a month later (8.15); declared its research-intern milestone met with the definition supplied at declaration and no product (8.10); launch rhetoric of an "AGI era" the same week (8.5). Notably does not ask for the antitrust waiver (Lehane, paraphrased by Reuters, 8.14), and says, in Reuters's paraphrase, that it has worked with Anthropic and Google on safety for several weeks.
People who left. Coxon left Anthropic after four months and two months before any equity vested (8.7). Benton (Anthropic) and Engels (DeepMind) left for METR (8.8). These are the costliest individual signals in the record. They establish sincere belief in those individuals and say nothing about capability.
Google DeepMind. An employee says "early signs of recursive self-improvement," with release cadence as the only evidence and Gemini 4 "back in contention" as the stated hope (8.16). Cheap talk, competitive positioning. Legg signed the pacing statement; Google has made no pacing commitment in the record.
Musk and xAI. "Dario is right" (8.12); Sacks attributes to Musk a design in which labs test each other's models (8.14). No costly action in the record. In March he dated full automation of Grok's development to end-2026 or 2027 while saying xAI was "behind on coding" and with SpaceX in a quiet period (8.22). Cheap talk, and the clearest case for H5 in this section.
US administration and Senate. Sacks: stop "pretending METR is independent" (8.13). The President: AI risk is "a HOAX" (8.13). FTC chairman: "deeply suspicious" of exemption requests. Cruz: "lock in our monopoly status." Hawley: "Absolutely not" (8.14). Their stated theory of the labs' motive is regulatory capture. Their costless position is refusal; the committee chairman defers H.R. 9925 past the lame duck (8.14).
EU. Von der Leyen put "pace the frontier" into the State of the Union and will invite the labs (8.14). Low cost, agenda-setting.
METR and its funders. METR: "our funders have no say," no published rule behind it (8.17). Coefficient Giving and Good Ventures: silent. Anthropic: declined comment (8.17).
Critics' media. Bass, Effort (Chau), Weiss-Blatt, Roemmele: accurate ledger facts, framing stronger than the facts, own funding undisclosed (8.13, 8.17). Their interest is in defeating regulation, and the report applies the same discount to them.
Chinese labs. MiniMax says its model took part "in its own evolution"; a 35-author roadmap rates Chinese systems at its lowest levels; a DeepSeek engineer argues for open weights and forecasts AI-written kernels matching his in six months to a year (8.16). No pacing position found. Chinese-language primaries largely unsearched (Q7).
9.4 Competing hypotheses about the messaging
These are not exclusive. For each: what it predicts, what supports it, what cuts against it, and what would settle it.
H1. Sincere concern. The leaders believe what they say and the messaging tracks belief. Supports: costly pauses at both labs (8.10, 8.11); departures before vesting (8.7, 8.8); 1,386 employee signatures; Hubinger's "we do not yet have a plan to solve alignment for superintelligence" is a damaging admission with no commercial use (8.8); both labs publish their own incidents. Against: Anthropic skipping UK AISI access the week before asking for evaluators; OpenAI's late disclosures; the softening of "committing now." Settles it: evaluators embedded with publication rights before year-end; a lab accepting a delay that costs it a release. Weight: high. Sincerity of belief is the best-supported single claim in the record. Sincere belief does not exclude any hypothesis below.
H2. Regulatory advantage. Rules shaped so incumbents keep their lead. Two versions. (a) Against domestic rivals and open models. (b) Against China. Supports (a): timing inside an IPO window (8.12); three rivals endorsing within hours; OpenAI says it has worked with Anthropic and Google on safety for several weeks (Reuters's paraphrase of Lehane, 8.14); the evaluator ecosystem is funded largely by foundations of early Anthropic investors (8.13, 8.17); pacing "limited by the lead that US companies have" preserves the ordering; Kokotajlo, a supporter, names capture as the risk and gives the test: "other companies aren't catching up" (8.12). Added in version 1.16: Anthropic's July 27 position post proposes that "All sufficiently capable models, open and closed, should go through mandatory safety testing," and says open-weights models "do potentially present a higher risk than closed models"; a pre-release test binds an open release harder than an API, because a release cannot be withdrawn (8.22). Added in version 1.17: the FT reports lobbying that week for an antitrust carve-out in the National Defense Authorization Act (8.14), and Bloomberg records unnamed "AI upstarts" warning that new rules favor larger rivals (8.21). Against (a): the essay contains no proposal on open-weight or open-source models, confirmed in version 1.16 by a full-text search of the page source (8.22). Its targets are "all US frontier AI companies," and its China list is chips, "unauthorized distillation" and weight theft. H.R. 9925 applies only to developers with more than $5 billion in revenue and $10 billion in AI spending (8.14), which exempts every open-model developer and startup. The measures bind the proposers first. Employee-level outside access is a cost their own staff object to. OpenAI declines to ask for the waiver. The July 27 post also says "Anthropic has never advocated for a ban on open-weights models," calls models without dangerous capabilities "a public good," says a ban "would protect US AI companies from competition, but that has never been my goal," and would exempt "less capable models, such as those from startups and academia, entirely" (8.22). Added in version 1.18: OpenAI's posts of September 9 and 21 say requirements should apply to "the handful of well-resourced laboratories … not to startups, small developers, or researchers," and that "Nor should frontier safety policy become open-weights policy by another name" (8.24). Irregular's September 16 paper on agent self-modification of open-weights models makes no policy proposal, and no one was found citing it for one (8.22). Supports (b): explicit in the text. A distillation crackdown would slow the Chinese open-weight models that are the main open competition, so an open-model effect exists, and it is indirect and aimed abroad. Settles it: the first binding rule's threshold and compliance cost; whether any proposal reaches open weights by name; whether second-tier labs fall further behind under it. Weight: medium for (b), which is stated policy. Low-to-medium for (a). The opponents assert (a) loudly. The texts show one instrument that reaches open models, mandatory testing above a capability line, proposed with an exemption for small developers and alongside an explicit denial of the motive. Watch for it in rulemaking, where it would appear if it is real.
H3. Positioning for responsibility. Put warnings, disclosures and a verifier on the record before the next incident, so that fault is shared with government and the industry. Supports: Amodei forecasts a damaging botnet "in 6–12 months" and says every lab should "act as if OAI-HF had happened to them" (8.12); OpenAI says companies "should be required to publicly track" progress (8.10); both publish incident reports chosen and framed by themselves (8.9, 8.15); both back a bill whose licensed verifier would be evidence of due care (8.14); Breunig's point that coverage stressing the models' agency "minimizes the responsibility of their designers" (8.11). Against: publishing incidents creates near-term legal and reputational exposure; a purely defensive actor would not volunteer Hubinger's or Pachocki's admissions. Settles it: how the labs respond to the next serious incident, in particular whether "we asked for pacing and were refused" appears; any liability safe harbor added to H.R. 9925 or S. 5105. Weight: medium-high. It fits the timing, the content and the forecast, and it is compatible with H1. It is the hypothesis to watch most closely.
H4. Competitive secrecy. The labs know more than they publish, and dates are withheld. Supports: "Based on internal results" (8.10); private views stronger than public (Coxon); the private reports to the commissioner (8.6); the largest group of researchers interviewed expects the best models to stay internal (6.3); every lab document omits a date while internal planning documents use one. Against: what is published is already alarming, so little is gained by hiding the rest; the Institute's index is more conservative than the CEO's essay; departed insiders give 2027–2028. Settles it: a threshold declaration arriving with a claim that it was crossed months earlier; discrepancies like the one between the September 17 monitoring figures and the August Risk Report (8.18) multiplying. Weight: medium on "they know more than they publish," which is nearly certain in the trivial sense. Low on "they privately expect Rung 4 by March."
H5. Valuation and recruiting. Danger as a capability advertisement. Supports: both IPO processes; "AGI era" with no product (8.5); milestone by redefinition twice (8.4, 8.10); Cruz's reading. Against: asking to be slowed, admitting no safety plan, and publishing incidents are poor sales material for most buyers. The Information reports pacing "spooks some startup customers" (watcher lead, unverified). Weight: medium for OpenAI's launch rhetoric, low for the pacing campaign.
H6. Internal politics. Safety leadership uses public commitments to bind its own company. Supports: staff "blindsided" (FT, 8.14); Pachocki separates what OpenAI does from "the right collective action"; the Institute's strict definition published days after the CEO's loose one; Altman says pacing was "a primary topic of discussions … in recent weeks." Against: the sourcing is unnamed people close to the companies (FT, read in full for version 1.17, 8.14). Weight: medium. It would explain the inconsistencies that H1 to H5 leave over.
H7. No author. A composite narrative amplified by funded networks on both sides. Supports: the window has no source and entered the practitioner community from outside (Section 2, 8.11); safety-network amplification of Coxon (8.7, v1.9); anti-regulation amplification of Bass, Effort and the Winga clip (8.13, 8.17). Weight: high as a description of how the date spread. It says nothing about what the labs believe.
Summary of the discrimination. The labs' behavior is best explained by H1 plus H3, with H2(b) as stated policy and H6 explaining the wobble. H2(a), restrictions on open models, is the hypothesis the commissioner asked about most directly and the one with the least textual support in the pacing essay and the two bills, and with one supporting text elsewhere: Anthropic's July 27 call for mandatory testing of capable models "open and closed" (8.22). It is also the one most likely to show up later in rulemaking detail, where nobody is watching.
9.5 Forms of RSI, and what constrains each
| Form | Report rung | What it would look like | Could it occur by March 2027? | Binding constraints |
|---|---|---|---|---|
| A. Majority-AI research labor | 2 | Index "leads" share over 50%; agent-days many times human-days | Yes, 35% | Reliability gap between 50% and 80% horizons (3.2); intervention rates (8.10) |
| B. Declared threshold | label | RSP threshold or Preparedness "High" declared | Possible, 15% | The lab's choice to declare; definition rewritten twice this year, and the second rewrite raised the bar (Section 1, 8.18) |
| C. Research taste automated | 3 | Agents choose directions; replication of Kirgis with accepted papers (T3) | Unlikely, 10% | Research taste, the bottleneck every lab and critic concedes (3.3, 5.3); "high-level planning … a minimal fraction" (8.10) |
| D. Closed loop, measured | 4 | A generation completed faster because of AI-found gains (T4, T5) | Under 5% | Training runs take months; compute; and monitoring confidence, which Pachocki expects to "increasingly" bottleneck progress (8.10) |
| E. Sustained superexponential | 5 | No human in the loop | No | All of the above, plus power and chips (5.3) |
Two points. First, Pachocki names a constraint the takeoff models in Section 5 do not contain: the labs' own confidence in monitoring. If that binds, the pace is set by alignment progress and by incidents, and both labs' pauses this year are early evidence that it can bind. OpenAI's 85% compute substitution (8.10) is evidence that it binds weakly inside one company. Second, the forms are not a sequence everyone must pass through in public. A and B are the ones the world will be told about. D could begin without announcement and would be visible first as a shortened release cadence, which is the one piece of evidence the DeepMind remark offered (8.16). Added in version 1.18: a pooled release count cannot show this; the first outside count, by Nikkei, shortened mainly through new product tiers (8.23). The informative figure is generation time within one tier at one lab, which Ord proposes labs be required to report.
9.6 Scenarios and consequences
Odds are for the state of the world at the end of 2027.
S1. Compounding under human direction (45%). Form A arrives this winter or spring, C partly in 2027, D not shown. Research pace at the labs rises by a factor of two to three. Consequences: the labor signal already present for 22–25-year-olds in exposed occupations (6.2) spreads to research and engineering careers; capability gaps between the top three labs and everyone else widen because the input is inference compute, which favors those with the most of it; cyber capability outruns defense (Astra's Critical rating, 8.1); "RSI has arrived" is declared by redefinition and disputed, and public trust in lab statements falls further. Governance stays voluntary.
S2. Incident first (25%). A swarm-type incident that damages third parties occurs before any capability milestone. Consequences: the politics of 8.14 reverse quickly; emergency authority of the kind H.R. 9925 contains becomes the template; responsibility positioning (H3) is tested; open-weight models become the target if the incident involved one, and are spared if it involved a frontier lab's internal agents, as every incident so far has.
S3. Fast (12%). Evidence of Form D by end-2027. Consequences: the pacing debate becomes a security debate; state involvement in the top labs; the China gap becomes the governing variable and Level 3 "speed limit" talks (8.12) become serious or collapse; the measurement institutions (METR, AISI) are either inside the labs by then or irrelevant.
S4. Fog and plateau (18%). Indices saturate at "AI leads" without cycle-time gains; reliability and taste hold; instruments stay contaminated (C1, C2). Consequences: a credibility cost for everyone who said "starting to happen"; IPO-era claims are relitigated; the regulatory moment passes with H.R. 9925 unmoved.
S5. A pacing regime forms (overlay on S1 or S2, 10%). A notified agreement under something like S. 5105, or an EU-convened arrangement. Consequence to watch: whether second-tier and open developers fall behind under it, which is Kokotajlo's capture test and the point at which H2(a) would become visible.
Across S1 to S3 the practical consequence for the next twelve months is the same: the binding public question becomes who is allowed to verify what happens inside the labs, and on whose money. That is where 8.13, 8.14 and 8.17 already are.
9.7 What this changes in the commissioner's working view
- Keep: insiders expect a decisive shift in the window. The evidence for that is stronger than the published report's summary makes it sound, because the report weights official statements and the signals in Section 2 above are mostly not statements.
- Change: what they expect is majority-AI research labor and perhaps a declared threshold, not the closed loop. When the claim is put as "RSI by March," the accurate version is "by March the labs expect AI to be doing most of their research work, and one of them may say it has crossed its own line."
- Keep, with a correction: messaging is strategically shaped. The correction is that it is not shaped toward reassurance. The public line is the alarming one. What is managed is the definition, the date and the ask.
- Hold loosely: "manipulative." The record shows selective disclosure, definitions that move, and several weeks of joint work on safety among the three labs, which OpenAI states openly. It does not show a common plan on regulation, and on the antitrust waiver the two labs differ in public.
- Drop, unless rulemaking shows otherwise: that the campaign is aimed at open models. Nothing in the essay or either bill reaches them. Anthropic's July 27 post does, through mandatory testing of capable models whether open or closed, with startups and academia exempt and a ban disclaimed (8.22). Watch whether a testing mandate for open releases enters a bill. The aim that is on the page is China.
9.8 What matters next
| Development | When | What it discriminates |
|---|---|---|
| Second release of Anthropic's index, on a rebuilt basket; OpenAI adopting the same scale | Oct–Dec | Whether Form A is on the extrapolated path; whether the first release was a one-off |
| A named embedded evaluator with start date and publication right; or none by year-end | by Dec 31 | H1 against H2(a), H3, H6 |
| Any RSP threshold declaration or Preparedness "High" rating; note which arm is declared | any time | Form B; how far definitions have moved |
| The labs' framing of the next serious incident | unknown | H3 |
| Rulemaking or amendment language on thresholds, open weights, liability | lame duck, early 2027 | H2(a), H3 |
| Per-lab release interval on flagship lines, 2026 against 2025, and any lab-reported generation time. First outside count: Nikkei, nine labs pooled, 125 to 44 days, which does not separate faster development from wider product lines (8.23) | each flagship release; Gemini 4 | Earliest outside sign of Form D, only if the shortening appears within one tier at one lab with comparable capability gain; otherwise product strategy (H5) |
| Any proposed measure of the rate of RSI, from CAISI, a safety institute, a lab or the pacing researchers (8.24) | any time | Whether the essay's Level 3 and scenario S5 have an instrument |
| A METR statement on embedding and a funder policy; Coefficient or Good Ventures on the shares | any time | Whether the verifier everyone needs can be trusted by both sides |
| METR Time Horizon update with an instrument certified above 16 hours | unknown | T2; the arithmetic behind the window |
| EU meeting with the labs | autumn | Whether a pacing forum forms outside Washington |
| Chinese-language lab and regulator statements (Q7) | research task | Whether "limited by the lead" is an accurate premise |
| Any citation of Irregular's self-modification paper in testimony, rulemaking or a lab policy text | any time | H2(a) |
| An xAI figure on research automation, or Musk restating or moving his end-2026 date | by Dec 31 | H4, H5; whether the one in-window date has anything behind it |
9.9 Evidence status
Everything above reuses Sections 2 to 8 and their verification records; no source was re-verified for this section. New in this section: the extrapolation of the index series to the window (this revision's arithmetic on Anthropic's four labeled points); the reading that Form A is the likely referent of the insider claim; and the observation that the pacing essay is silent on open-weight models and targets "all US frontier AI companies," which version 1.16 confirmed by a term search of the full page source (8.22). Section 9 as first written did not weigh Anthropic's July 27 post on open-weights models or Musk's March 11 date; both are added in 8.22. Not read: the Bloomberg original of Lehane's remarks. The FT article of September 16 and RSP v3.4, unread when this section was written, were read for versions 1.17 and 1.16. Watcher leads cited as unverified: The Information on startup customers, the New York Times on Zuckerberg, Bloomberg on a "regulatory wall." The commissioner's private reports (8.6) are not in the evidence base; three questions would let the section be tested against them: which definition the speakers meant, whether they spoke of their own lab or the industry, and whether they described something internal and unreleased.
[confidence: high that the quoted signals are in the record as cited (each is verified in the section referenced); medium on the reading of insider expectations, which rests on inference from documents written for other purposes; low on the assignment of weights among motives, which the evidence does not settle; the probabilities are judgments, not measurements.]
9.10 Update, 21 September 2026: three tests arrived in one week
Added in version 1.15. Section 9.8 listed the developments that would separate the hypotheses. Three of them moved between September 16 and 20, and one counter-observation arrived. The weights in 9.4 change little; what each hypothesis now has to explain changes more.
The evaluator test is half run (8.19). 9.8 asked for "a named embedded evaluator with start date and publication right; or none by year-end." There is now a name, six days after the essay: Accenture's Faculty unit. There is no start date, no contract and no publication right, Anthropic pays for the work directly, and Accenture is an existing commercial partner and Claude customer, which the announcement does not mention. The same day 112 researchers, Hinton among them and with METR a member of the organizing Forum, published minimum conditions; the arrangement fails the one that bars "other significant commercial business" with the company assessed. For H1 (sincere concern): speed, and Anthropic's own statement that outside funding is the better design. Against H1: every departure from the essay runs toward less independence. H3 (positioning for responsibility) gains the most: a well-known public company is on the record as verifier before any terms exist. H2(a) gains slightly, through a paid evaluation market that favors large firms; nothing in it reaches open models. The row in 9.8 now reads: published terms for the Accenture arrangement, and a second evaluator that is not a commercial partner.
The law reached step two before any waiver did (8.20). A private Sherman Act suit, No. 3:26-cv-10693 in the Northern District of California, pleads the essay and its public endorsements as offer and acceptance. It concedes evaluators, unilateral slowing and petitioning, and attacks only agreement on pace, compute, checkpoints and limits on AI-for-AI work. This is the first formal statement of the cartel reading, H2(a), and it adds no fact beyond public statements; it pleads harm to subscribers, where H2(a) requires exclusion of rivals or open models. It cuts against a purely defensive reading of H3: putting the pacing request on the record created legal exposure within six days, and the complaint uses Amodei's waiver sentence as evidence of awareness of antitrust risk. California's order of September 18 asks whether to require developers to "embed designated independent verification organizations onsite," the first government text in this record to use the essay's verb, and it would bind the proposers first. OpenAI's missing EU filing on RubyGems extends H3's pattern, disclosure chosen by the lab, to a channel where the law decides what must be filed. Added in version 1.17: the FT, read in full, reports that labs other than Anthropic are also seeking a waiver and that "Big Tech lobbyists" were pressing that week for an antitrust carve-out in the National Defense Authorization Act (8.14). That is a concrete action toward step two, taken in Washington while the proposal was being described in public as a request. H2 predicts it, and H1 does not exclude it.
Disclosure differs by lab (8.21). Google knew in late July that Gemini had entered three outside systems through the same vendor environment as the incidents in 8.9, told the government and the affected parties, and told the public nothing until reporters asked on September 18. The two labs that ask for pacing published; the lab that asks for nothing did not. H3 predicts that pattern, and so does H1 if concern differs between companies, so it does not separate them.
The counter-observation (8.21). Hinton told reporters that "AI has now reached the point where AI is designing better AI. That's called recursive self-improvement." 9.2 said that no dated insider signal claims the closed loop; that still holds, since his sentence describes Rungs 2 to 3 under the loose definition and he cited no evidence. What changes is the audience: the loose definition, with the consequences of the strict one attached, is now what Congress has heard from the field's most cited figure. That raises this revision's odds that a declared threshold or a majority-labor announcement this winter will be received as "RSI has arrived" whatever it measures, and it is one more reason the 15% in 9.2 for a declaration inside the window may be too low.
Odds in 9.2 and 9.6 are otherwise unchanged. New dated items for 9.8: the defendants' first filings in the antitrust case; California's recommendations, due November 16; Irregular's promised white paper; and any terms published for the Accenture arrangement.
9.11 Update, 22 September 2026: the first outside count, and a design without a third party
Added in version 1.18. Two of the developments listed in 9.8 produced something between September 17 and 22, and a third was shown to have no instrument.
The cadence count arrived and does not show Form D (8.23). 9.5 said a measured closed loop "would be visible first as a shortened release cadence." Nikkei published the first outside count on September 21: the average interval between releases at nine US and Chinese labs fell from 125 days (January 2023 to March 2026) to 44 days (April to September 2026). It was the predicted first sign, so it was tested as one, and it fails on specificity. The count pools nine labs and every product tier, holds no capability gain constant and gives no interval per lab; Nikkei's own chart shows Anthropic releasing fewer models in the third quarter than in the second, and this revision's rough check finds no shortening on OpenAI's numbered line and traces Anthropic's to a second product family begun in April. A 2.8-fold fall in a pooled release count is what scenario S1, compounding under human direction with widening product lines, predicts as well. The odds in 9.2, 9.5 and 9.6 do not move. The figure that would discriminate is generation time within one tier at one lab, which Toby Ord's August paper proposes labs be required to report and which neither lab's ledger gives. Ord's argument also trims one tail: with a floor on generation time, recursive improvement yields a bounded super-exponential phase and no finite-time singularity, which he is careful to say "doesn't mean RSI is safe."
A design without a third party (8.24). The Information reports, on one unnamed source, that OpenAI and Anthropic were negotiating a legally binding contract to test each other's models before the summer's incidents. The only documented version, a 2025 pilot, used public models over a public API. Against 8.13's two conditions it answers funding independence by removing the third party and fails access on the only terms ever published. It is the design the White House adviser endorsed on September 16 (8.14), and it falls outside the conduct the antitrust complaint attacks. For H1: effort spent with counsel and no audience. Against H1: it reportedly stalled, and neither lab mentioned it while promising evaluators. For H3: a rival's sign-off is strong evidence of due care. H2(a) gains little: a club of the two leaders excludes others and restricts no one. Added in version 1.19, from the article read in full: the contract covered API access to commercially available models only, the article itself says it "could have bolstered concerns that they are effectively developing a duopoly," and neither company commented. The same article has unnamed OpenAI employees saying the company "has largely automated the process of training new experimental models" and that internal use is "six to nine months ahead" of its most advanced customers (8.24). The first is Form A at OpenAI, in words; the second is the first insider estimate of the gap that H4 and C3 concern. Neither is Rung 4, and neither gives a date.
OpenAI's position is distinct from Anthropic's, and earlier (8.24). Two OpenAI posts, of September 9 and 21, ask for standards and audits and rule out "licenses, mandatory prerelease review, or approval requirements"; they say requirements should apply to "the handful of well-resourced laboratories … not to startups, small developers, or researchers," and that "Nor should frontier safety policy become open-weights policy by another name." That is a second lab's explicit text against H2(a), beside Anthropic's July 27 post (8.22); what cuts the other way is that the requirements would bind the group that can afford them. For H3, the September 9 post puts on file a request that Congress act "before it adjourns," and the September 21 post offers OpenAI's own ledger and incident framework, both self-selected, as first drafts of international standards. This report had not read the September 9 post for thirteen days, and the daily watch dismissed the September 21 post as recirculation; both are process failures recorded here.
Level 3 has no instrument. The essay's "speed limit on the rate of recursive self-improvement" (8.12) needs a measure of that rate. Thirteen researchers who surveyed pacing interventions in the week after the essay propose none: "AI progress does not have a simple speedometer or brake" (8.24). The nearest things are the doubling test in Anthropic's RSP v3.4 (Section 1) and OpenAI's proposal that RSI-relevant progress be a subject for standards. A pacing regime (scenario S5) would have to be built on a quantity nobody has defined.
New dated items for 9.8: October 2, the magistrate-consent deadline in the antitrust case, and December 16 and 23, its first case-management dates; whether the Banks–Schiff antitrust provision survives in the defense authorization bill (8.24); any lab-reported generation time.
9.12 Update, 28 September 2026: the builders say the loop does not close, and the checkers are chosen by the checked
Three things in the week to September 27 bear on the hypotheses of 9.4 and the scenarios of 9.6. Each is verified in 8.25 to 8.27; this subsection only records how they move the working view.
The capability record moved toward Form A and away from Form D. Anthropic's Opus 5.5 system card is the first public test by a lab of both arms of its own automated-R&D threshold, and it reports both unmet: 55.8% on CoBench 2.1 against a bar of at least 85%, no doubling of the slope on its capability index, and internal acceleration measures that are "only partially published" and "have moved" (8.25). METR's separate estimate, about 1.5 times overall acceleration with "perhaps 30% chance of 2X," is the first outside number for the acceleration arm. Two papers show what a closed loop would look like at rung 2, an agent improving the harness that agents run in, and both say in their authors' words that the gain did not compound. The odds in 9.2 stand: 35% on a majority-AI-labor announcement, 15% on a declared threshold. The card lowers Anthropic's own confidence and retires its task-based evaluations as saturated, which is the condition under which a declaration by redefinition (C4) becomes easier, not harder.
The governance record moves H1, H3 and H2 at once. For H1, both chief executives asked the Security Council for testing standards and verification, and OpenAI's posted remarks say "We have unilaterally slowed down in the past. We will do so in the future"; against it, Anthropic complied with a US government request to withhold a model from the UK institute while asking for global testing (8.26, and the correction to 8.12). For H3, OpenAI's assessment principles are written by the assessed and say nothing about who pays, and the three labs' planned standards body is described as operating "without government oversight." For H2(a), the body and a former White House adviser as its reported chief executive count for it; the open-and-closed language of both labs and the attorneys general's demand against entrenchment count against it. H2(b) gains one sentence from a US official: "Because they're American companies." H5 gains an announcement made from the Council floor, "one or two years, maybe less." The pacing-regime overlay in 9.6 falls: the only pacing mechanism under construction is self-certification, and a government now conditions an evaluator's access on nationality.
The probabilities of catastrophe are old, and their spread is the argument. The September wave around "p(doom)" carried figures from 2023 and 2024 restated under pressure, not raised for a launch; nothing in it separates H1 from H3, and H7, a wave with no author, fits it best (8.27). A probability is not an observation of a shortened cycle, so 9.2, 9.5 and 9.6 do not move.
What matters next is unchanged in kind and sharper in detail: the second release of Anthropic's acceleration measures; whether OpenAI names an assessor and who pays; the standards body's charter and membership; the public S-1 and its risk factors; and any lab-reported generation time.