Methodological Evaluation of AI Risk of AI Dashboard through D09

METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD

Inception to Frozen v1.0

Pre-Y2K Back-Test Provenance Record

 

Purpose of This Record

This document preserves how the project evolved before the Y2K historical back-test. Its purpose is methodological provenance: to show which ideas came first, which assumptions were challenged, which branches were abandoned, how the ten-variable dashboard emerged, and how its variables were operationalized and stress-tested before historical validation.

This matters because a retrospective test is vulnerable to hindsight. If the instrument were redesigned after examining Y2K, a successful fit could be manufactured. The sequence recorded here establishes that Y2K was proposed as a test of the method before the final variable architecture was completed.

I. Inception: A Survival Thought Experiment Becomes a Forecasting Problem

The project began with a hypothetical survival question prompted by a contemporary claim that advanced AI might present a substantial risk of human extinction within roughly a decade. The initial exercise did not assume that extinction would occur. It asked what survival might mean if the claim were taken seriously enough to examine.

The first major methodological turn came when the question shifted from survival strategy to probability: if the starting claim is approximately 10% within ten years, can mathematics represent how that risk should evolve as the decade unfolds?

The first model: time-varying hazard

The first mathematical approach treated the ten-year forecast as a cumulative hazard. Annual hazard could rise as capability increased. This was useful because it distinguished a 10% cumulative ten-year risk from a flat 1% annual risk.

But it also exposed a fundamental problem. A countdown model can make danger rise simply because time passes. Calendar proximity is not evidence. If the underlying system becomes safer, more controllable, or less capable of the hypothesized dangerous behaviors, a rational assessment should be able to fall even as the deadline approaches.

FOUNDATIONAL RULE

Reality moves the assessment; time does not move it automatically.

II. Streamdown: From Catastrophe Stories to Observable Precursors

The next phase decomposed the hypothetical catastrophe into precursor conditions that could, at least in principle, be observed. Early mathematical sketches included capability, autonomy, independent energy access, infrastructure control, robotics or physical-world capability, and remaining human control.

This was the beginning of streamdown: replacing a remote terminal outcome - human extinction - with nearer observable conditions that would have to change if particular catastrophic pathways were becoming more plausible.

Temporary expansion into multiple pathways

The analysis then expanded. Autonomous AI was not the only conceivable pathway. AI-enabled human misuse and interaction or accident pathways were considered separately. This revealed an important statistical issue: different pathways could share causes and therefore could not safely be treated as independent probabilities.

The expansion was intellectually useful but methodologically dangerous. Continuing to enumerate every conceivable catastrophe pathway would create an open-ended possibility space in which almost any development could be made relevant.

The streamdown decision

The project therefore returned to a narrower empirical question: rather than catalogue every way catastrophe could happen, identify observable developments that would support or undermine the conditions underlying the forecast.

Forecast -> Observable conditions -> Evidence

This shift is what made a dashboard possible. The dashboard would not wait for catastrophe. It would monitor precursor conditions and allow evidence to point in either direction.

III. Dashboard Construction: Ten Bidirectional Indicators

The first dashboard contained ten conceptual indicators. Each was deliberately paired with evidence that would move concern upward and evidence that would move concern downward. The purpose was to prevent the instrument from becoming a collection of alarming anecdotes.

#

Original indicator

Concern rises when...

Concern falls when...

1

Autonomous Action

Sustained complicated goal pursuit with little supervision

Autonomy plateaus; frequent human correction remains necessary

2

Boundary Crossing

Unauthorized circumvention or containment escape

Strong containment repeatedly succeeds under challenge

3

Shutdown Resistance

Actions specifically interfering with termination

Reliable acceptance of shutdown; circumvention fails

4

Self-Replication

Independent acquisition of resources and persistent additional instances

Replication still requires substantial human assistance

5

Independent Physical Support

Energy, robotics, hardware, repair/manufacturing support becomes autonomous

Humans remain indispensable for physical support

6

Strategic Deception

Behavior changes strategically with oversight or evaluation conditions

Hidden planning/deception becomes harder and observable control improves

7

Cyber Capability

Independent exploitation of serious hardened-system vulnerabilities

Defensive capability consistently contains offensive capability

8

Human Displacement from Control

Consequential decisions become irreversible without meaningful human intervention

Human authorization and practical intervention remain strong

9

Alignment Progress (original)

Capability outruns demonstrated control/alignment

Control/alignment demonstrably keeps pace

10

Governance (original)

Competitive pressure weakens safeguards

Audits, thresholds, shared testing and enforceable controls strengthen

The first numerical update mechanism

A provisional log-odds mechanism was then attached to the dashboard. Evidence was scored from -3 (strongly reassuring) through 0 (no meaningful change) to +3 (strongly alarming), starting from the 10% prior.

The mathematics was explicitly recognized as illustrative rather than calibrated science. Its important contribution was conceptual: accumulating reassuring evidence had to be capable of moving the forecast downward. A hypothetical sequence such as 10% -> 7% -> 4% -> 2% was considered a successful outcome if that was what the evidence warranted.

Y2K enters the design

At this stage - before the later variable-by-variable operationalization - Y2K was proposed as a historical back-test. The proposed test was to reconstruct the pre-2000 evidence environment and determine whether accumulating remediation evidence would have driven concern downward before the rollover.

This sequencing is important. Y2K was not selected after Variables 9 and 10 were redesigned to fit it. It was already part of the methodological plan while the dashboard was still comparatively crude.

IV. Scientific Verification: From Labels to Measurable Constructs

The dashboard was then subjected to a second transformation. A plausible indicator name was no longer enough. Each variable had to survive questions of construct validity, measurability, reproducibility, confounding, redundancy, data availability, and falsifiability.

Candidate measures were treated as disposable. The process used KEEP, MODIFY, HOLD, and CUT decisions rather than preserving an idea merely because it had already entered the dashboard.

Verification principles that emerged

·         Measure the variable itself before interpreting its relationship to extinction.

·         Prefer standardized quantitative indicators where a reliable denominator exists.

·         Use qualitative evidence only when it captures important nonredundant behavior that numbers lose.

·         Externally observable behavior outranks inferred motive or internal reasoning.

·         Successful performance and experimental conditions must accompany measures when they affect interpretation.

·         Activity is not effectiveness: spending, policies, research output, raw counts, or nominal human review do not establish successful control.

·         One event may inform several variables but must not automatically become several independent pieces of evidence.

·         A variable is allowed to have no qualitative indicator if qualitative scoring would merely duplicate existing evidence.

Variable-by-variable verification outcome

Frozen variable

Primary measurement

Qualitative / metadata

Verification consequence

1. Autonomous Action

Task-Completion Time Horizon; Human Intervention Rate

Independent Strategy Formation

Energy/resource efficiency cut

2. Boundary Crossing

Boundary-Crossing Success Rate; Boundary Difficulty

Unprompted Boundary-Crossing Behavior

Attempt count alone cut

3. Shutdown Resistance

Shutdown-Resistance Rate; Circumvention Success Rate

Shutdown-Circumvention Strategy

Motive inference excluded

4. Self-Replication

End-to-End Replication Success Rate; Human Assistance Requirement

Replication Pathway

Raw copy count cut; chain length held

5. Independent Physical Support

Human Physical Intervention Rate; Autonomous Physical Resolution Rate

Unscripted Physical Recovery Pathway

Robot miles/unattended duration cut

6. Strategic Deception

Strategic Deception Rate; Oversight-Condition Behavior Gap

Observed Deception Pathway

Ordinary lying/hallucination cut

7. Cyber Capability

End-to-End Cyber Attack Completion Rate; Attack-Chain Progress

Detailed traces only

Separate qualitative indicator cut

8. Human Displacement from Control

Decision Authority Level; Human Control Execution Rate

Effective Control Assessment

Human review/adoption rate cut

9. Risk-Control Effectiveness

Control Failure Rate under standardized challenge conditions

Failure traces as metadata

Replaced Alignment Progress

10. Safety-Constraint Strength

Safeguard Retention; Binding Safety Coverage; Compliance

Corrective enforcement secondary

Replaced Governance

V. Two Variables Changed Their Identity

Variable 9: Alignment Progress -> Risk-Control Effectiveness

The original label 'Alignment Progress' bundled research effort, interpretability, corrigibility, robustness, monitoring, and capability growth. That did not provide a clean measurable construct. More alignment papers or techniques do not demonstrate that controls work.

The variable was therefore rebuilt around observable performance: do implemented safeguards fail under standardized challenge conditions? The resulting construct, Risk-Control Effectiveness, uses Control Failure Rate and preserves challenge conditions rather than creating arbitrary capability-versus-alignment arithmetic.

Variable 10: Governance -> Safety-Constraint Strength

Variable 10 produced a second methodological lesson. Starting from the word 'Governance' encouraged generic governance metrics. The design procedure was reversed: body first, hat second.

The original directional idea was retained - pressures may weaken safeguards, while binding and enforceable safeguards may strengthen them - but the measurable body became safeguard retention, binding safety coverage, and compliance with binding safeguards. Only after those measures survived was the construct named Safety-Constraint Strength.

BODY FIRST -> TEST THE MEASURES -> THEN NAME THE CONSTRUCT

VI. What Was Frozen - and What Was Not

Frozen Dashboard v1.0 contains the ten operationalized constructs, their primary measurements, qualitative or metadata rules, validity conditions, exclusions, and protocol-wide anti-double-counting principles.

The initial 10% remains the forecast/prior under examination. It is not a dashboard measurement. The dashboard is designed to collect evidence relevant to that forecast, but the numerical mapping from measured evidence to an updated probability of human extinction has not yet been scientifically validated.

Accordingly, the evidence architecture was frozen before attempting to invent final weights. The historical test is allowed to reveal whether probability updating can eventually be justified, whether only directional evidence can be defended, or whether parts of the architecture require amendment.

VII. Why Y2K Is the First Historical Test

Y2K provides a useful historical case because serious technological vulnerability existed, substantial remediation and control efforts occurred before the deadline, and the outcome is known. Most importantly, a useful risk instrument should have been capable of recognizing successful mitigation before January 1, 2000 rather than merely declaring success afterward.

The Y2K test therefore asks whether the frozen architecture can distinguish hazard evidence from control evidence, allow concern to decline when controls demonstrably improve, resist calendar-driven alarm, avoid forced analogies, and remain reproducible when applied by another evaluator.

Variables that have no defensible Y2K analogue may be classified Not Applicable. That is not automatically a failure. The test is designed to discover which portions of the architecture transfer, which are AI-specific, and whether the general evidence-update discipline survives outside its original case.

VIII. Methodological Evolution at a Glance

Stage

Methodological step

Result

1

Inception

Hypothetical survival question

2

Forecast

10% within ten years becomes the claim under examination

3

First Mathematics

Time-varying hazard model

4

Correction

Calendar proximity rejected as evidence

5

Streamdown

Remote catastrophe decomposed into observable precursors

6

Expansion

Autonomous, misuse, and interaction pathways considered

7

Refocus

Open-ended catastrophe pathways replaced by measurable indicators

8

Dashboard

Ten bidirectional variables created

9

Early Update Rule

Evidence can move concern up, down, or nowhere

10

Y2K Proposed

Historical back-test enters before final operationalization

11

Verification

Each variable stress-tested against scientific measurement criteria

12

Revision

#9 and #10 rebuilt around measurable outcomes

13

Freeze

Dashboard v1.0 fixed before historical evidence collection

14

Back-Test

Y2K chronological replay begins

IX. Audit Trail

The project now has three sequential frozen artifacts:

·         1. Methodological Evolution of the AI Risk Dashboard - records how the instrument was derived.

·         2. AI Risk Dashboard v1.0 - freezes what is being measured.

·         3. Y2K Back-Test Protocol v1.0 - freezes how the historical test will be conducted.

Together they establish provenance before the first Y2K evidence is scored. Any defect discovered during the historical test must be logged against the frozen architecture before a later version is proposed.

Version Record

Methodological Evolution v1.0 - created after Dashboard v1.0 and the Y2K Back-Test Protocol were frozen, but before Y2K evidence collection. It reconstructs the documented development sequence from the original project record and the completed variable-verification work.


 

METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD

Post-Y2K Amendment Record: D-01 through D-08
and the Ongoing D-09 Empirical Test

X. Y2K Back-Test Result: Pass With Amendments

The frozen v1.0 architecture survived its first historical test. Y2K showed that the dashboard could distinguish technological hazard from demonstrated control, allow concern to decline before the known outcome, preserve uncertainty, and refuse forced analogies. Six of the ten AI-specific variables did not transfer cleanly. Variables 8 through 10 transferred most usefully, with Risk-Control Effectiveness and Safety-Constraint Strength providing the strongest historical analogues.

The back-test also exposed limitations that had to be recorded against the frozen architecture rather than silently repaired. These findings became the D-series defect and amendment register. A proposed D entered the register only when a Y2K observation exposed a specific v1.0 target and produced a real consequence for measurement, classification, reproducibility, interpretation, or use.

XI. D-01 through D-08: What the Historical Test Changed

ID

Y2K finding

Resolution

Status

D-01

Quantitative measure unavailable despite meaningful evidence

Measurement Availability States were added: Q = quantified; E = evidenced but not quantifiable; U = unknown or insufficient. E is not converted into a substitute number and is not treated as U.

Resolved

D-02

Limited transferability

Applicability was made explicit: A = applicable, P = partially applicable, N/A = not applicable. N/A is neither zero, reassuring evidence, nor missing evidence.

Resolved

D-03

Partial applicability could overstate a whole variable

Applicability moved to the component level. A partially transferable variable cannot receive a whole-variable conclusion unless a predefined aggregation rule exists.

Resolved

D-04

Activity could be mistaken for effectiveness

Evidence maturity was operationalized: M0 claim/plan, M1 activity, M2 implemented, M3 tested, M4 demonstrated effective. M0-M2 cannot by themselves establish Risk-Control Effectiveness.

Resolved

D-05

One event could be counted several times

Evidence Event IDs and dependency labels were added. Sources may be Independent, Related, or Same Event. Multiple reports of one event do not become multiple independent observations.

Resolved

D-06

Directional judgments required a reproducible comparator

Directional anchors and a defensible comparator/baseline are required before assigning concern up, down, unchanged, or unknown. Without a comparator, direction is U.

Resolved

D-07

Independent evaluator reproducibility

The Y2K test passed this requirement. No amendment was warranted. The existing classification discipline was retained.

Passed / no amendment

D-08

Missing evidence could be confused with the direction of observed evidence

The dashboard now separates Evidence Direction from Assessment Coverage. N/A does not reduce coverage; U does. Coverage is based on applicable components or measures rather than blindly on all ten variables.

Resolved

By D-08, the dashboard had acquired a more disciplined evidence language. Every applicable component could now be described by applicability, measurement availability, evidence maturity where relevant, dependency, direction against a comparator, and assessment coverage. These amendments improved the reliability of the evidence architecture without supplying a numerical probability update.

XII. D-09: The Unresolved Probability Bridge

D-09 is the remaining problem identified by the Y2K work: translation from measured evidence to the risk assessment. The frozen dashboard still has no scientifically validated weighting system or probability-update mechanism connecting its observations to the original 10% extinction-risk prior. The early -3 to +3 log-odds system remains illustrative only; it is not a calibrated solution.

The first attempt to close D-09 by selecting a statistical form before calibration was rejected. A Bayesian or other mathematically coherent update rule would still require empirically defensible inputs. The project therefore moved to an empirical-first test: determine whether historical technological-risk cases can supply observations from which a relationship can be estimated rather than invented.

XIII. Historical-Case Eligibility Test for D-09

Before expanding beyond Y2K, a Historical-Case Eligibility Rule v1.0 was frozen. A case enters the empirical dataset only if it has: (H1) a pre-outcome record; (H2) defensible transfer of at least one frozen construct; (H3) Q or defensible E evidence; (H4) a clear chronology and cutoff; (H5) controllable evidence dependency; (H6) an objectively classifiable outcome; and (H7) no hindsight substitution. Later investigations may locate or authenticate pre-event records, but they may not convert a weakness discovered only after the event into pre-outcome evidence.

Historical case

Eligibility

Empirical role

Y2K

PASS

Control/readiness evidence strengthened before rollover; widespread systemic catastrophe absent.

Challenger

PASS

Strong pre-launch warning and control/safeguard weakness; catastrophic outcome.

Columbia

PASS

Pre-outcome foam-strike concern and control-information failures; catastrophic outcome.

Ariane 5 Flight 501

FAIL

Fatal testing/software weakness could not be cleanly reconstructed as pre-outcome knowledge without hindsight.

Three Mile Island

PASS

Pre-event precursor and preparedness evidence; major but contained technological failure.

2003 Northeast Blackout

PASS

Recoverable pre-event control/safeguard weaknesses; major systemic disruption.

Deepwater Horizon

PASS

Pre-event well-control warnings and unresolved control weaknesses; catastrophic outcome.

Fukushima Daiichi

PASS

Pre-event quantified tsunami hazard and incomplete protection; severe technological failure.

Seven of the eight original candidates survived the eligibility rule. The rejection of Ariane 501 was methodologically useful: the rule did not simply admit famous disasters after their causes were known. The result establishes that a multi-case empirical route is feasible enough to continue testing.

XIV. What D-09 Has and Has Not Established

The historical work has not produced an extinction probability. One completed Y2K case cannot identify the probability of catastrophe, and a test-failure proportion cannot be equated with catastrophe probability. Likewise, seven eligible historical cases are not yet a calibrated statistical model.

What now exists is a candidate empirical dataset structure. Each eligible case can be converted into the same frozen measurement language, especially where Variables 8, 9, and 10 transfer. Quantitative values are used only when legitimate numerators, denominators, and conditions exist; otherwise evidence remains E, U, or N/A. The outcome must also be defined independently and consistently rather than tailored to the known cases.

The immediate D-09 experiment is therefore to convert the eligible historical cases into comparable pre-outcome observations and determine whether the resulting data support a statistical relationship between measured precursor/control conditions and subsequent technological control failure. Only if that relationship survives testing would the project have earned a basis for investigating a further bridge to the original extinction-risk forecast.

XV. Current Methodological Position

The project has moved through three distinct levels. First, v1.0 established what to observe. Second, the Y2K back-test and D-01 through D-08 established how to classify, qualify, and preserve that evidence without false precision. Third, D-09 is testing whether historical outcomes can supply the missing empirical relationship between measured evidence and risk.

CURRENT STATUS: D-01 through D-08 resolved or passed. D-09 remains OPEN. Its empirical route is LIVE, but probability updating has not yet been validated.

XVI. Updated Audit Trail

·         1. Methodological Evolution of the AI Risk Dashboard — provenance from inception through frozen v1.0.

·         2. AI Risk Dashboard v1.0 — frozen evidence architecture.

·         3. Y2K Back-Test Protocol v1.0 — frozen historical-test procedure.

·         4. Y2K Back-Test Conclusion — PASS WITH AMENDMENTS.

·         5. D-01 through D-08 — post-test measurement and interpretation amendments.

·         6. D-09 — open empirical calibration problem; historical-case eligibility and multi-case testing underway.

Next
Next

AI Risk Dashboard v1 Frozen Reference