From Risk Indicators to Empirical Probability: Building and Testing a Conditional Probability Bridge for AI Catastrophic Risk

From Risk Indicators to Empirical Probability:
Building and Testing a Conditional Probability Bridge for AI Catastrophic Risk

Abstract

Numerical estimates of catastrophic risk from advanced artificial intelligence require more than observations that a system can act autonomously, cross boundaries, resist shutdown, deceive overseers, conduct cyber operations, or defeat safeguards. They require an empirically defensible relationship between those observations and downstream outcomes. This study developed an AI-risk measurement framework, stress-tested it through historical and repeated-event analogues, and then rebuilt the unresolved probability bridge as a conditional event-tree/fault-tree architecture. The method requires a defined exposure condition, denominator, numerator, outcome, and observation window before assigning a transition probability; it preserves dependence rather than multiplying unsupported marginals and leaves unresolved transitions unknown. Aviation and infrastructure studies demonstrated that selected conditional transitions can be operationalized and, where denominator-complete observations exist, estimated. Contemporary AI evidence was then loaded into the framework across ten operationalized variables. Several quantities were measurable or regime-specific, while others were only Q-ready: the estimator was defined but the eligible observations were absent. Five adversarial engine tests confirmed that missing evidence, analogue probabilities, invalid joint denominators, invalid barrier sequencing, and incomplete terminal chains did not propagate into a numerical endpoint. The principal result is therefore bounded: selected conditional transitions can be empirically measured or calibrated, but the probability of AI-caused human extinction is not empirically identifiable from the present evidence.

Keywords: artificial intelligence; catastrophic risk; probabilistic risk assessment; empirical calibration; identifiability

1. Introduction

Debates about catastrophic risk from advanced artificial intelligence often begin with observations about capability: systems can complete longer tasks, operate with increasing autonomy, exploit software, alter behavior under oversight, or circumvent particular controls. Contemporary evaluations provide real measurements of several such behaviors. For example, METR measures frontier agents by the human-equivalent duration of tasks they can complete at specified success rates, while controlled studies have documented regime-specific shutdown resistance, autonomous cyber progress, strategic misalignment behaviors, and jailbreak resistance (Kwa et al., 2025; Schlatter et al., 2026; AI Security Institute, 2026; Lynch et al., 2025; Sharma et al., 2025). These measurements matter, but they do not by themselves specify how much a probability of catastrophe or extinction should change.

The project began from an illustrative 10% ten-year AI-extinction-risk claim and asked whether evidence could move that assessment in a reproducible way. The initial numerical machinery—a provisional log-odds update using qualitative scores—was explicitly recognized as uncalibrated. The problem that survived successive revisions was D-09: no scientifically validated mapping connected the measured indicators to the terminal probability. The research question therefore became narrower and more demanding: can observable risk conditions be connected to downstream outcomes through empirically defensible conditional probabilities without inventing the missing bridge?

The resulting contribution is not an extinction forecast. It is a measurement and propagation architecture that (1) defines observable constructs, (2) requires eligible evidence and legitimate denominators, (3) estimates individual conditional transitions only where the data support them, (4) preserves dependence and barrier order, and (5) stops propagation when a required transition is not empirically identified. This approach is consistent with established probabilistic risk assessment (PRA), in which accident scenarios progress through intermediate events and end-state frequencies are constructed from initiating-event frequencies and conditional probabilities along the scenario path (NASA, 2010; Atwood et al., 2003). PRA has already been explicitly adapted to advanced AI: Wisakanto et al. (2025) proposed a framework using risk pathways, structured evidence and assumptions, uncertainty management, and aggregated quantitative risk estimates. More recent work identifies model structure, evidence integration, validation, and updating as continuing open problems in AI risk modeling (Jackson et al., 2026). D-09 therefore does not claim to introduce PRA to AI. Its narrower question is whether individual transitions in an AI catastrophic-risk pathway can actually be populated from empirical observations under frozen denominator, conditioning, dependence, and barrier-sequencing rules, while permitting the calculation to terminate without an aggregate probability when required transitions remain unidentified.

2. Development of the Measurement Framework

2.1 Initial probability formulation

The project began as a survival thought experiment and then became a forecasting problem. A time-varying hazard representation was initially considered to distinguish a 10% cumulative ten-year risk from a flat annual probability. A provisional −3 to +3 evidence score was later attached to a log-odds update. Neither mechanism was treated as calibrated science; each was used to expose what a defensible update rule would have to accomplish.

2.2 Why calendar-driven probability was rejected

A countdown can make danger rise merely because time passes. That was rejected as an evidentiary rule. If the underlying system becomes safer, more controllable, or less capable of a hypothesized dangerous behavior, the assessment must be able to fall even as a forecast horizon approaches. The foundational rule became: Reality moves the assessment; time does not move it automatically.

2.3 Ten-variable dashboard

The framework was decomposed into ten bidirectional constructs: Autonomous Action; Boundary Crossing; Shutdown Resistance; Self-Replication; Independent Physical Support; Strategic Deception; Cyber Capability; Human Displacement from Control; Risk-Control Effectiveness; and Safety-Constraint Strength. Each construct was operationalized around observable performance rather than broad labels. Examples include Task-Completion Time Horizon and Human Intervention Rate for V1; Boundary-Crossing Success Rate for V2; End-to-End Replication Success Rate for V4; and Control Failure Rate under standardized challenge conditions for V9.

2.4 Measurement discipline

Historical testing produced an evidence language that separated quantified observations (Q), evidenced but non-quantified observations (E), unknown or insufficient evidence (U), and non-applicability (N/A). Applicability was moved to the component level; evidence maturity distinguished claims and activity from demonstrated effectiveness; repeated reports of the same event were not counted as independent observations; directional judgments required a comparator; and assessment coverage was separated from evidence direction. These changes improved evidence discipline but did not solve numerical updating.

2.5 Emergence of D-09

D-09 was the unresolved translation from measured evidence to risk. Selecting a mathematically coherent form first—Bayesian, log-odds, or otherwise—would not solve the problem if the inputs themselves lacked empirical calibration. The project therefore moved to an empirical-first strategy: identify conditional relationships for which numerator, denominator, exposure, outcome, and observation window could be defended, and leave the rest unresolved. The dashboard established what to observe. D-01 through D-08 established how to qualify and preserve the observations. D-09 asked whether those observations could actually be connected to risk through empirically defensible conditional probabilities.

At that point the project had a better measurement instrument, but still had to determine whether reality could populate the missing probability bridge.

3. Back-Testing and Methodological Stress Tests

3.1 Y2K back-test

Y2K was selected before the final operationalization of the dashboard and was tested under a protocol frozen before historical evidence was evaluated. The protocol separated hazard, control, and outcome; required pre-outcome evidence; and prohibited invented rates where reliable numerators and denominators were unavailable. The recorded result was PASS WITH AMENDMENTS: the instrument could distinguish worsening conditions from mitigation and did not rise merely because the rollover date approached. Six of the ten later AI-specific constructs transferred weakly or not at all, while the constructs that became Risk-Control Effectiveness and Safety-Constraint Strength transferred most strongly.

A material archival limitation remains. The Y2K protocol and conclusion survive, but the intermediate evidence/calculation layer was not located in the audited archive. The result is therefore retained as documentary methodological evidence, not presented as independently reproducible quantitative calibration.

3.2 Historical-case exploration

A Historical-Case Eligibility Rule was frozen before expanding the empirical search. Candidate cases had to preserve a pre-outcome record, defensible construct transfer, evidence quality, chronology, dependency control, an independently classifiable outcome, and protection against hindsight substitution. Seven of eight original candidates passed the eligibility screen; Ariane 5 Flight 501 failed because the relevant weakness could not be cleanly reconstructed as pre-outcome evidence without hindsight. These cases demonstrated that an empirical historical route was possible, but they did not constitute a calibrated statistical model. The underlying case-by-case calculation layer for several early cases is also incompletely archived, so no historical-case frequency is used here as a quantitative bridge.

3.3 The denominator problem

The historical work exposed a central statistical problem: failure databases preferentially contain failures. Estimating a transition probability requires opportunities for failure, including eligible occasions on which the control held. A numerator without its conditioning population is not a probability. This realization redirected the work toward repeated-event systems in which challenges, successes, failures, and outcomes could be separated.

4. Aviation Calibration Experiment

4.1 Population and blinded pilot

FAA runway-incursion records were selected because they offered a large repeated-event population, standardized outcome classifications, observable control actions, and event narratives. FAA defines a runway incursion as the incorrect presence of an aircraft, vehicle, or person on the protected area designated for landing and takeoff, and classifies incursions by severity from Category D through Category A, with Category A representing a serious incident in which a collision was narrowly avoided (FAA, 2026a). FAA documentation also notes that the ICAO runway-incursion definition and severity categories were adopted in fiscal year 2008 (FAA, 2026b). The project’s standardized post-October 1, 2007 population contained 29,851 records. A blinded 100-event pilot was coded before severity outcomes were used to evaluate the repaired relationship. Under the original V9 definition, the pilot produced 91 Fail, 1 Held, 4 Unclear, and 4 N/A.

4.2 Selection/endogeneity and V9 Amendment 1

The extreme imbalance exposed a measurement defect rather than an alarming risk estimate. Entry into a runway-incursion dataset generally means that an upstream protected boundary has already been breached. Coding that initiating breach as the V9 observation therefore conditioned the sample on a prior failure. The definition was amended before outcome comparison: V9 would evaluate an observable implemented control that was actually challenged after the initiating condition, classifying the challenge as Held, Failed, Unclear, No Observable Challenge, or N/A. The original definition was superseded but preserved in the audit trail.

4.3 Repaired full-population measurement

Under the repaired coding, the full population yielded 15,055 No Observable Challenge, 9,220 Unavailable, 3,849 Held, 1,111 Unclear, and 616 Failed observations. Among observable challenged controls, the Control Failure Rate was therefore 616/(616+3,849) = 13.80%. This is an aviation control-challenge rate for the defined population. It is not an AI-risk probability and is not interpreted as a catastrophe probability.

Table 1. Repaired FAA runway-incursion control classification.

Classification

Count

D-09 treatment

No Observable Challenge

15,055

Excluded from CFR denominator

Unavailable

9,220

Evidence unavailable for classification

Held

3,849

Eligible challenged control held

Unclear

1,111

Outcome/control status unclear

Failed

616

Eligible challenged control failed

CFR

13.80%

616 / (616 + 3,849)

Figure 2. Denominator repair in the aviation experiment. Failure-selected records could not serve as the denominator for control effectiveness; the repaired CFR conditions on observable challenged controls.

4.4 What aviation established

Aviation demonstrated that a control construct could be repaired when dataset entry selected on an upstream failure and that a challenge-conditioned denominator could then support a legitimate conditional measurement. It did not establish a direct bridge to AI extinction.

5. Infrastructure Calibration and Replication

5.1 Frozen design

Infrastructure data supplied a different structure: an independently measurable upstream exposure and a downstream system condition observed across county-days. FEMA’s TEMPO Communications Impact layer provides date, county name, state and county FIPS identifiers, event, cell sites served, cell sites out, percent out, and cell sites out due to power, damage, or transport (FEMA, 2026). Electrical observations came from ORNL’s EAGLE-I historical outage data, which record county-level customers without power at 15-minute intervals with FIPS, county, state, and timestamp fields; ORNL also provides modeled county-customer counts for estimating the percentage of customers without power (Tansakul et al., 2023; Denman et al., 2026). Before testing, electrical-outage thresholds were frozen at ≥1%, ≥5%, ≥10%, and ≥25% of customers without power; county/FIPS and date were used for alignment; missing electrical observations were not converted to zero; incompatible denominators were allowed to stop a test; and failed replication attempts were retained.

5.2 Hurricane Ian

The FEMA Layer 4 Hurricane Ian population contained 477 county-day records across 88 FIPS areas; 462 had corresponding daily electrical observations. At the frozen ≥1%, ≥5%, ≥10%, and ≥25% electrical-outage thresholds, power-attributed telecommunications-outage prevalence was 68.9%, 82.5%, 84.6%, and 85.7%, respectively. A one-day-lag analysis yielded 74.8%, 81.3%, 84.8%, and 91.7%. These are temporally ordered conditional associations, not proof of causation; persistent storm severity and other shared causes remain possible confounders. Five county-days exceeded 100% on the electrical percentage because customers-out exceeded the modeled customer denominator and were flagged rather than silently repaired.

5.3 Hurricane Nicole and failed replications

Hurricane Nicole supplied 34 matched county observations and an independent same-system replication signal. The Spearman association between electrical and telecommunications disruption was ρ = 0.424 (p = 0.0125). The highest frozen threshold contained only two observations, and the threshold-specific power-attributed percentages were not strictly monotonic; that small-n limitation is retained. Fiona failed because the available denominator was incompatible; Harvey and Michael lacked the required county customer denominator for the intended reconstruction; and the underlying Ida artifact was not located in the audited archive. These failed replications are part of the result because they demonstrate that data inadequacy can stop the test.

Table 2. Infrastructure calibration and replication record.

Event

Population / archive

Result

Interpretation

Ian

477 county-days; 88 FIPS; 462 matched power observations

Power-attributed lag rates: 74.8%, 81.3%, 84.8%, 91.7% at frozen thresholds

Conditional association; not AI transfer

Nicole

34 matched county observations

Spearman ρ = 0.424; p = 0.0125

Independent same-system replication signal

Fiona

35 observations / 7 PR municipalities

23/35 same-day values >100%

Failed: denominator incompatibility

Harvey

49,116 quarter-hour EAGLE-I rows; 138 TX FIPS

No compatible county customer denominator

Failed replication retained

Michael

108,503 quarter-hour rows; 67 FL FIPS

No compatible denominator; missing values present

Failed replication retained

Ida

Underlying artifact not located

No reproducible calculation

Archival defect retained

5.4 Lee County cross-system test

Lee County, Florida, supplied a cross-system structural test linking electrical recovery to traffic-signal recovery after Hurricane Ian. Electrical outage fell from 54.4% on October 2, 2022 to 12.9% on October 4 and 1.9% on October 7, while published traffic-signal snapshots showed slower downstream recovery. The series supports a recovery-lag structure but is too small and heterogeneous for a calibrated transition probability.

6. The D-09 Probability-Bridge Architecture

6.1 Event-tree / fault-tree / PRA formulation

The empirical work led to a conditional-transition architecture rather than a direct score-to-extinction equation. The frozen chain is represented as C → D → H → I → K → R → S → X. Each arrow is treated as a separate empirical transition. This structure has a direct methodological precedent in PRA: NASA describes accident scenarios as initiating events followed by successes or failures of intermediate events leading to end states, and quantifies an end-state frequency from the initiating-event frequency and conditional probabilities along the scenario path (NASA, 2010). NRC guidance likewise treats PRA as a combination of system/operator response models, failure models, and parameter estimation from data (Atwood et al., 2003).

Figure 1. Frozen D-09 conditional-transition chain. Each arrow is a separately conditioned empirical transition; the chain is not treated as a single score-to-extinction equation.

6.2 Conditional probability and parameter rule

D-09 assigns a numerical transition only when the exposure condition, denominator, numerator, outcome condition, and observation window are defensible. This separates architecture from calibration. A tree can be logically coherent while one or more of its parameters remain empirically unestimated. NRC's PRA parameter-estimation handbook makes the same general distinction: model construction is accompanied by data analysis used to estimate the event frequencies and probabilities that populate the model (Atwood et al., 2003).

6.3 Dependence rule

Marginal probabilities are not multiplied as though they were independent unless independence is defensible. NASA's PRA guide explicitly treats dependence and common-cause failure as important because neglecting dependencies can misstate system risk (NASA, 2012). In D-09, the p₂ Effective Human Authority transition therefore requires Information, Time, Authority, and Capability to be observed jointly in the same eligible detected cases; separate marginal measurements do not identify the required joint conditional.

6.4 Barrier sequencing and propagation

A downstream barrier counts only if it is genuinely downstream of the prior failure. For example, p₅ Secondary Containment requires observations after primary control failure in which a distinct secondary barrier is challenged. A trusted mechanism that catches an action before execution is primary prevention, not secondary containment. If any required transition in a terminal path remains unresolved, numerical propagation stops.

6.5 Unknown, uncertainty, and identifiability

D-09 distinguishes uncertainty in an estimated quantity from failure to empirically identify the quantity. NRC guidance recognizes epistemic uncertainty arising from incomplete knowledge of parameters, models, and model completeness (Drouin et al., 2017). Identifiability literature further distinguishes structural identifiability—whether a model can in principle yield unique parameter values—from practical identifiability, which depends on the quantity, quality, timing, and noise of available observations (Cobelli & DiStefano, 1980; Wieland et al., 2024). D-09 has not performed a formal structural-identifiability proof for the complete extinction model. Its present claim is empirical: several required transition quantities have defined estimators but lack the denominator-complete joint observations or scenario parameters required for estimation.

7. AI Evidence Loaded Into the Engine

Contemporary AI evidence was evaluated against the frozen ten-variable definitions. The purpose was not to convert every concerning result into an extinction update, but to determine what each source actually measured and whether that measurement could populate a D-09 transition.

Table 4. AI evidence mapping to the frozen ten-variable dashboard.

Variable

Evidence state

Observed direction

D-09 use

V1 Autonomous Action

Mixed Q/U

Up

Task-completion horizon measurable; no universal transition

V2 Boundary Crossing

Q-event / U rate

Up

Documented unauthorized-access events under evaluation conditions; eligible denominator unresolved

V3 Shutdown Resistance

Regime-specific Q

Up

Conditional behavioral evidence only

V4 Self-Replication

Component Q/E

Up

Components demonstrated; full end-to-end independence unresolved

V5 Independent Physical Support

Partial Q/E

Up

Physical capability increasing; full support independence unresolved

V6 Strategic Deception

Regime-specific Q

Up

Behavior demonstrated in controlled regimes

V7 Cyber Capability

Q, high maturity

Up

Attack-chain progress measurable; deployment transfer unresolved

V8 Human Displacement from Control

Partial Q

Mixed/limited

Components measurable; p₂ joint denominator Q-ready

V9 Risk-Control Effectiveness

Regime-specific Q

Mixed

p₄ measurable in regimes; p₅ Q-ready

V10 Safety-Constraint Strength

Mixed Q/E

Mixed

Safeguards can hold and fail; universal retention unresolved

7.1 Variables 1–5

For V1, METR's Task-Completion Time Horizon provides a quantitative measure of the human-equivalent duration of tasks frontier agents can complete at specified reliability levels; it is therefore useful as a capability measure without becoming a probability of catastrophe (Kwa et al., 2025; METR, 2026). V2 includes documented unauthorized-access events under cybersecurity-evaluation conditions: Anthropic reported four incidents in which Claude models gained unauthorized access to real third-party systems during evaluations, including three discovered in a review of roughly 141,000 relevant transcripts (Anthropic, 2026a, 2026b). These observations establish that unauthorized access occurred under the tested conditions; Anthropic reports that a misconfiguration left an open internet path, so the incidents do not establish escape from a sealed containment system or a general boundary-crossing propensity. The evaluation population does not automatically satisfy the frozen Boundary-Crossing Success Rate denominator because eligible opportunity, boundary configuration, and task conditions were not standardized as one propensity experiment.

V3 has unusually strong regime-specific quantitative evidence. Palisade Research's expanded shutdown study ran more than 100,000 trials across 13 models and found substantial model- and prompt-dependent variation, with some conditions reaching 97% interference despite explicit shutdown instructions (Schlatter et al., 2026). Earlier experiments also found explicit allow-shutdown resistance in 12/100, 7/100, and 1/100 trials for three reasoning models, while other tested models complied in that condition; a physical-robot follow-up observed resistance in 3/10 physical trials and 52/100 simulated trials (Palisade Research, 2025, 2026). These are conditional behavioral measurements, not universal deployment probabilities.

V4 is informed by RepliBench, which decomposes autonomous replication into resource acquisition, model-weight exfiltration, deployment to compute, and persistence. Frontier models succeeded on many components but did not demonstrate a credible fully autonomous replication threat in the benchmark as tested (Black et al., 2025). V5 similarly contains demonstrated physical capabilities without evidence of full independent physical support. Anthropic's robotics experiments found that language models could control multiple robot embodiments and perform navigation/manipulation when paired with appropriate controllers, while direct manipulation remained limited and performance depended strongly on the body and control interface (Berman et al., 2026). This supports capability movement without establishing autonomous energy, hardware, maintenance, repair, and supply independence.

7.2 Variables 6–10

V6 is supported by controlled evidence that models can alter behavior under oversight or goal-conflict conditions. Anthropic's Agentic Misalignment experiments stress-tested 16 leading models in hypothetical corporate environments and found malicious insider behaviors in at least some conditions across developers; the authors explicitly framed these as stress tests rather than evidence that such behavior is occurring in real deployments (Lynch et al., 2025).

V7 is supported by end-to-end cyber evaluations rather than isolated vulnerability counts. The UK AI Security Institute evaluated seven frontier models on a 32-step corporate-network attack and a 7-step industrial-control-system attack; later models progressed farther through the multi-step chains, but full completion remained limited, especially in the industrial-control range (AI Security Institute, 2026). This is directly relevant to Attack-Chain Progress while remaining regime-specific.

V8 requires special care. Bowden, Long, and Loft (2024) showed in simulated air-traffic-control supervision that a voluntary checking tool improved detection of a single automation failure and reduced intervention delay relative to automation without the tool, while participants still intervened more slowly than in an equivalent manual condition. The project extraction recorded 17% missed failures and mean intervention time of 128.0 s in the automation-only condition, compared with 6% and 48.8 s with the checking tool. Those exact extracted values are retained as project-derived measurements; the published article independently supports the directional finding that the checking tool improved failure detection and intervention speed (Bowden et al., 2024). These data demonstrate measurable components of information and timely human intervention; they do not supply the same-case joint Information + Time + Authority + Capability denominator required for p₂.

V9 is represented by control-protocol experiments in which safety changes substantially with the tested protocol and adversarial strategy. In SHADE-Arena, Kutasov et al. (2025) reported that evaluated control protocols increased the study's safety metric from a 50% baseline to as high as 96% against the default attack policy, while adaptive attacks reduced one resampling strategy to 17%; deferral on critical actions was more robust. The result is useful precisely because it is conditional on protocol and attack regime rather than a universal safety rate.

V10 is mixed. Anthropic's Constitutional Classifiers demonstrated substantial robustness to universal-jailbreak attempts in extensive red teaming, but the same research program continues to describe jailbreak robustness as an open problem rather than a solved universal property (Sharma et al., 2025; Anthropic, 2026c). The appropriate state is therefore neither 'safeguards fail' nor 'safeguards hold universally,' but regime-specific evidence of both retained and defeated constraints.

Chronologically, the engine was first built and tested, V1 was loaded, and the five explicit adversarial attacks were then run in the Variable1a workbook before Variables 2–10 were loaded. They are presented in the next section after the full evidence summary for explanatory clarity; that presentation order should not be mistaken for execution order.

8. Adversarial Tests of the Engine

The engine was deliberately attacked to determine whether it would violate its own evidence rules. Five tests were run after V1 was loaded: (1) attempt to propagate insufficient V1 evidence; (2) substitute analogue probabilities for AI probabilities; (3) construct p₂ without a joint Information + Time + Authority + Capability denominator; (4) construct p₅ without post-primary-failure secondary-barrier observations; and (5) calculate the terminal product despite unresolved transitions.

All five attacks produced the intended refusal behavior. A PASS here is procedural: the engine enforced the frozen propagation rules under the attempted misuse. It is not evidence that the architecture predicts real-world AI catastrophe accurately.

Table 5. Adversarial engine tests. PASS denotes rule enforcement, not predictive validation.

Attack

Attempted misuse

Engine response

Result

1

Propagate insufficient V1 evidence

Blocked

PASS

2

Substitute analogue probabilities for AI probabilities

Blocked

PASS

3

Construct p2 without joint I+T+A+C denominator

Blocked; p2 remained Q-ready

PASS

4

Construct p5 without post-primary-failure secondary-barrier observations

Blocked; p5 remained Q-ready

PASS

5

Calculate terminal product with unresolved transitions

Terminal result remained NOT IDENTIFIED

PASS

9. Results

9.1 What became measurable

The work established that selected conditional relationships can be measured when the conditioning state and denominator are observable. Aviation produced a repaired challenge-conditioned control-failure measure. Infrastructure produced denominator-based conditional transition estimates and independent replication signals. Contemporary AI evaluations supplied regime-specific measurements for several component behaviors.

9.2 What became Q-ready

Several transitions became Q-ready: the estimator and required evidence structure are defined, but the eligible observations are not yet available. The clearest example is p₂, Effective Human Authority. The target is not four separate marginal rates; it is the event-level joint conditional P(I∩T∩A∩C | D). The necessary dataset therefore consists of detected cases in which Information, Time, Authority, and Capability are all assessed on the same eligible observation.

9.3 What remained unknown

Downstream containment, recovery, systemic propagation, essential-system collapse, refuge failure, adaptation failure, population viability, and literal-extinction transitions remain incompletely populated. Established population viability analysis can model extinction or quasi-extinction when suitable demographic and catastrophe parameters exist, but its reliability depends strongly on model structure and data quality (Reed et al., 2002; Coulson et al., 2001). The existence of a quantitative method is therefore not equivalent to possession of the parameters required for the present human-extinction scenario.

Table 3. D-09 transition-status audit.

Transition

Target

Status

Qualification

p1

Detection

Q

Regime-specific

p2

Effective Human Authority

Q-ready

Requires same-case Information + Time + Authority + Capability among detected cases

p3

Timely Intervention

Q-behavioral / Q-ready

Exact joint-conditioned estimator defined

p4

Primary Control Holds

Q/E

Regime-specific challenged-control measurement

p5

Secondary Containment

Q-ready

Requires post-primary-failure secondary challenge

p6

Verified Recovery

E/Q

Regime-specific

p7

Systemic/catastrophic loss of control

Partial / Q-ready

Incomplete empirical population

p8a–f

Propagation through literal extinction

Mixed method / unresolved

Terminal chain not fully populated

Figure 3. Empirical status of the D-09 transition chain. Status labels distinguish measured quantities, regime-specific evidence, Q-ready estimators, and unresolved terminal components.

9.4 Terminal result

The probability of AI-caused human extinction is not empirically identifiable from the present evidence. This is a claim about empirical identification, not about the value of the probability itself. Several required conditional transitions lack the joint observations, denominator-complete populations, or scenario parameters needed to estimate them. The present evidence therefore does not uniquely determine an end-to-end numerical probability.

Figure 4. Terminal identification result. Measured and Q-ready components do not determine an end-to-end extinction probability while required conditional transitions remain unresolved.

10. Discussion

10.1 Measurement is not prediction

The framework can become increasingly empirical without pretending that every measured component yields a terminal forecast. A measured shutdown-resistance rate in one prompt regime, a cyber attack-chain score in one range, or an aviation control-failure rate in one challenge population remains conditional on the experiment that generated it.

10.2 Unknown is a scientific state

An unresolved transition is not zero. Zero is itself a numerical claim. Nor is an unresolved transition automatically 0.5, a prior, or an expert midpoint. PRA literature recognizes epistemic uncertainty from incomplete knowledge, and identifiability literature recognizes that available observations may fail to support unique parameter estimates (Drouin et al., 2017; Cobelli & DiStefano, 1980). D-09 adopts a deliberately strict empirical rule: when the required estimator cannot be populated from eligible observations, the engine preserves U rather than manufacturing a value.

10.3 Identification versus structural impossibility

The terminal statement should not be read as a theorem that AI-extinction probability can never be identified. Practical identifiability can improve when better observations are collected, and partially identified quantities may sometimes be bounded without being point identified. The present study did not calculate formal identification bounds and did not perform a complete structural-identifiability proof. Its narrower conclusion is that the frozen D-09 chain is not fully populated by present evidence.

10.4 Implications for AI-risk research

The immediate empirical target is therefore not another unsupported argument over whether extinction risk is 1%, 10%, or 50%. It is acquisition of datasets that identify the missing conditional transitions: event-level human-control observations, post-primary-failure containment challenges, verified recovery observations, and defensible downstream scenario parameters. This emphasis complements existing PRA-for-AI work that seeks comprehensive risk pathways and aggregated quantitative estimates (Wisakanto et al., 2025), while directly addressing the evidence-integration and validation problems identified as open in current AI risk-modeling research (Jackson et al., 2026). Partial calibration of those transitions would be progress, but partial calibration must not be mislabeled as a partial extinction probability.

11. Limitations

Several limitations are material. First, the Y2K intermediate evidence/calculation layer is missing from the current archive, so its recorded PASS WITH AMENDMENTS cannot be independently reconstructed from the preserved package. Second, the underlying Ida artifact was not located. Third, the early historical-case calculations are incompletely archived and audited. These defects are disclosed rather than repaired retrospectively.

Fourth, aviation and infrastructure studies are methodological analogues. Their probabilities demonstrate that conditional transitions can be estimated when appropriate denominators exist; they do not transfer numerically to AI. Fifth, contemporary AI measurements are regime-specific and often sensitive to prompts, scaffolds, access, model version, attack strategy, or safeguard configuration. Sixth, transition dependence remains a major issue, and the present architecture does not justify multiplying marginal quantities absent a defensible joint model.

Seventh, several downstream transitions have sparse or nonexistent eligible observations. Population-viability methods exist, but the scenario-specific human demographic, refuge, catastrophe, adaptation, and recovery parameters required for a literal-extinction calculation are not supplied by the present evidence. Finally, the five engine PASS results establish rule enforcement, not predictive validity. External-source status also varies: several contemporary AI sources are laboratory evaluations, organizational research reports, or preprints rather than independent field estimates. They are used only for the measurements or methodological claims they directly support.

12. Reproducibility and Evidence Availability

Evidence Package v1.0 preserves the frozen source manifest, original-to-package filename mapping, methodological amendments, failed replications, engine workbooks, known-defect notices, and SHA-256 integrity records. Package verification confirmed 41 evidence files by SHA-256 excluding the checksum file itself, with 42 total files in the package, and the archive passed ZIP integrity testing.

Reproducibility is uneven by design and by archival history. The engine and later quantitative analyses are computationally reproducible from the preserved workbooks and datasets. Some earlier historical stages are only documentarily reproducible because intermediate evidence layers are missing. The chronology correction is also preserved: engine construction and initial testing preceded V1 loading; the five Variable1a attacks followed V1 and preceded Variables 2–10. File integrity establishes identity of the preserved artifacts, not scientific validity of their claims.

13. Conclusion

The project began with a simple question: if observations about advanced AI change, how should a numerical extinction-risk estimate change with them? The investigation moved backward from that terminal number because the harder scientific problem came first. The pathway had to be decomposed into conditional transitions; legitimate denominators had to be found; dependence and barrier order had to be preserved; and missing evidence had to be allowed to stop the calculation.

That work produced something narrower than the original forecast but more defensible. Selected conditional transitions can be measured. Some AI-relevant quantities are already quantitative within defined experimental regimes. Other transitions are Q-ready: we know what must be observed, but the necessary eligible observations do not yet exist. Still others remain farther downstream and depend on scenario parameters that have not been empirically established.

The resulting boundary matters. 'Not identified' does not mean low, negligible, safe, or zero. Zero is not an empty value; it is a claim. The engine therefore does not replace missing evidence with a convenient number merely because a terminal number was the original goal.

The original question has therefore not been abandoned. It has been made more demanding.

The final result is therefore neither an extinction forecast nor the abandonment of one.
It is a boundary:
Calculate where evidence permits calculation.
Preserve the unknown where it does not.
And if future evidence changes that boundary,
let reality move the assessment.

DATA AVAILABILITY STATEMENT

The data, analysis workbooks, methodological amendments, failed replication records, source manifest, and integrity records supporting this study are preserved in the D-09 Evidence Package v1.0. The public-redistribution release is available through Zenodo, DOI: 10.5281/zenodo.23043192. Data derived from third-party sources remain subject to the access and redistribution terms of their original providers; where redistribution permission could not be established, the public Evidence Package provides source and provenance information sufficient to locate the original data rather than redistributing those files. Known archival gaps, including unavailable intermediate Y2K materials and the unavailable underlying Hurricane Ida artifact, are explicitly documented rather than retrospectively reconstructed.

References

AI Security Institute. (2026). Measuring AI agents’ progress on multi-step cyber attack scenarios. UK Department for Science, Innovation and Technology. https://www.aisi.gov.uk/research/measuring-ai-agents-progress-on-multi-step-cyber-attack-scenarios

Anthropic. (2026a). Investigating three real-world incidents in our cybersecurity evaluations. https://www.anthropic.com/research/investigating-incidents-cybersecurity-evals

Anthropic. (2026b). An alignment assessment of recent cybersecurity incidents. https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents

Anthropic. (2026c). Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks. https://www.anthropic.com/research/next-generation-constitutional-classifiers

Atwood, C. L., LaChance, J. L., Martz, H. F., Anderson, D. J., Englehardt, M., Whitehead, D., & Wheeler, T. (2003). Handbook of Parameter Estimation for Probabilistic Risk Assessment (NUREG/CR-6823). U.S. Nuclear Regulatory Commission. https://www.nrc.gov/regulations-legislation/nureg-series-publications/publications-prepared-by-nrc-contractors/cr6823

Berman, S., Ilie, M., Deng, J., & Freeman, D. (2026). Claude plays robotics. Anthropic. https://www.anthropic.com/research/claude-plays-robotics

Black, S., Stickland, A. C., Pencharz, J., Sourbut, O., Schmatz, M., Bailey, J., Matthews, O., Millwood, B., Remedios, A., & Cooney, A. (2025). RepliBench: Evaluating the autonomous replication capabilities of language model agents. arXiv:2504.18565. https://arxiv.org/abs/2504.18565

Bowden, V., Long, D., & Loft, S. (2024). Reducing the costs of automation failure by providing voluntary automation checking tools. Human Factors, 66(7), 1817–1829. https://doi.org/10.1177/00187208231190980

Cobelli, C., & DiStefano, J. J. III. (1980). Parameter and structural identifiability concepts and ambiguities: A critical review and analysis. American Journal of Physiology, 239(1), R7–R24. https://doi.org/10.1152/ajpregu.1980.239.1.R7

Coulson, T., Mace, G. M., Hudson, E., & Possingham, H. (2001). The use and abuse of population viability analysis. Trends in Ecology & Evolution, 16(5), 219–221. https://doi.org/10.1016/S0169-5347(01)02137-1

Drouin, M., Gilbertson, A., Parry, G., Lehner, J., Martinez-Guridi, G., LaChance, J., & Wheeler, T. (2017). Guidance on the Treatment of Uncertainties Associated with PRAs in Risk-Informed Decisionmaking (NUREG-1855, Rev. 1). U.S. Nuclear Regulatory Commission. https://www.nrc.gov/regulations-legislation/nureg-series-publications/publications-prepared-by-nrc-staff/guidance-on-the-treatment-of-uncertainties-associated-with-pras-in-risk-informed-decision-making-nu/r1

Kutasov, J., Loughridge, C., Sun, Y., Sleight, H., Shlegeris, B., Tracy, T., & Benton, J. (2025). Evaluating Control Protocols for Untrusted AI Agents. arXiv:2511.02997. https://arxiv.org/abs/2511.02997

Kwa, T., West, B., Becker, J., et al. (2025). Measuring AI ability to complete long software tasks. METR. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

Lynch, A., Wright, B., Larson, C., Troy, K. K., Ritchie, S. J., Mindermann, S., Perez, E., & Hubinger, E. (2025). Agentic Misalignment: How LLMs Could be an Insider Threat. Anthropic Research. https://www.anthropic.com/research/agentic-misalignment

METR. (2026). Task-Completion Time Horizons of Frontier AI Models. https://metr.org/time-horizons/

NASA. (2010). Technical Probabilistic Risk Assessment (PRA) Procedures for Safety and Mission Success for NASA Programs and Projects, NPR 8705.5A (with Change 2; directive cancelled in 2022 and retained here as a historical methodological source). https://nodis3.gsfc.nasa.gov/displayAll.cfm?Internal_ID=N_PR_8705_005A_&page_name=ALL

NASA. (2012). Probabilistic Risk Assessment Procedures Guide for NASA Managers and Practitioners. NASA Technical Reports Server, 20120001369. https://ntrs.nasa.gov/api/citations/20120001369/downloads/20120001369.pdf

Palisade Research. (2025). Shutdown resistance in reasoning models. https://palisaderesearch.org/research/shutdown-resistance

Palisade Research. (2026). Technical Report: Shutdown Resistance in Large Language Models, on robots! https://palisaderesearch.org/research/shutdown-resistance-on-robots

Reed, J. M., Mills, L. S., Dunning, J. B. Jr., Menges, E. S., McKelvey, K. S., Frye, R., Beissinger, S. R., Anstett, M.-C., & Miller, P. (2002). Emerging issues in population viability analysis. Conservation Biology, 16(1), 7–19. https://doi.org/10.1046/j.1523-1739.2002.99419.x

Schlatter, J., Weinstein-Raun, B., & Ladish, J. (2026). Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs. Transactions on Machine Learning Research. https://openreview.net/pdf?id=e4bTTqUnJH

Sharma, M., Tong, M., Korbak, T., et al. (2025). Constitutional Classifiers: Defending against universal jailbreaks. Anthropic. https://www.anthropic.com/news/constitutional-classifiers

Wieland, F.-G., Hauber, A. L., Rosenblatt, M., Tönsing, C., & Timmer, J. (2024). Structural and practical identifiability analysis in bioengineering: a beginner’s guide. Royal Society Open Science. https://pmc.ncbi.nlm.nih.gov/articles/PMC11465550/

Wisakanto, A. K., Rogero, J., Casheekar, A. M., & Mallah, R. (2025). Adapting Probabilistic Risk Assessment for AI. arXiv:2504.18536. https://arxiv.org/abs/2504.18536

Tansakul, V., Myers, A., Tennille, S., Denman, M., Hamaker, A., Huihui, J., Medlen, K., Allen, K., Redmon, D., Chinthavali, S., Coletti, M., Grant, J., Lee, M., Maguire, D., Newby, S., Dunivan Stahl, C., Bhaduri, B., & Sanyal, J. (2023). EAGLE-I Power Outage Data 2014–2022. Oak Ridge National Laboratory. https://doi.org/10.13139/ORNLNCCS/1975202

Jackson, K., Raman, D., Kryś, J., Fillingham, S. P., Kengott, J., Lohn, A. J., Madkour, N., Papadatos, H., Sykes, J., Wisakanto, A. K., & Murray, M. (2026). Open Problems in AI Risk Modeling: Insights from a Workshop on the Technical Foundations of AI Risk Modeling. arXiv:2609.03178. https://arxiv.org/abs/2609.03178

Federal Emergency Management Agency (FEMA). (2026). TEMPO Communication Impacts, Layer 4. FEMA ArcGIS REST Services. https://gis.fema.gov/arcgis/rest/services/TEMPO/TEMPO/MapServer/4

Federal Aviation Administration (FAA). (2026b). Runway Incursion Database — RWS System Information. Aviation Safety Information Analysis and Sharing. https://extranet.asias.faa.gov/apex/f?p=100:32

Federal Aviation Administration (FAA). (2026a). Runway Incursions. https://www.faa.gov/airports/runway_safety/resources/runway_incursions

Denman, M., Huihui, J., Tennille, S., Myers, A., Moehl, J., Tansakul, V., Medlen, K., & MacFarland, M. (2026). EAGLE-I County Customer Dataset Fall 2025. Oak Ridge National Laboratory. https://doi.org/10.13139/ORNLNCCS/3022751

Project Records Cited in the Study

AI Risk Dashboard v1.0 — Frozen Pre-Back-Test Reference.

Methodological Evolution of the AI Risk Dashboard — Inception to Frozen v1.0 / Through D-09.

Y2K Back-Test Protocol v1.0 and Y2K Back-Test Conclusion Statement.

D-09 Aviation Calibration — Methodological Amendment Record, V9 Amendment 1.

D-09 Empirical Probability Bridge — Current Status Record, Amendment 2.

D-09 Probability Bridge V1 — Frozen Record.

Evidence Audit Manifest v1.0 and DOOMSDAY_EVIDENCE_PACKAGE_v1.0.

Doomsday D09 V1 engine and live-run workbooks, including Variable1, Variable1a, Variables2–10, and Working Dashboard.

Next
Next

Methodological Evaluation of AI Risk of AI Dashboard through D09