Methodological Evaluation of AI Risk of AI Dashboard through D09
METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD
Inception to Frozen v1.0
Pre-Y2K Back-Test Provenance Record
Purpose of This Record
This document preserves how the project evolved before the Y2K historical back-test. Its purpose is methodological provenance: to show which ideas came first, which assumptions were challenged, which branches were abandoned, how the ten-variable dashboard emerged, and how its variables were operationalized and stress-tested before historical validation.
This matters because a retrospective test is vulnerable to hindsight. If the instrument were redesigned after examining Y2K, a successful fit could be manufactured. The sequence recorded here establishes that Y2K was proposed as a test of the method before the final variable architecture was completed.
I. Inception: A Survival Thought Experiment Becomes a Forecasting Problem
The project began with a hypothetical survival question prompted by a contemporary claim that advanced AI might present a substantial risk of human extinction within roughly a decade. The initial exercise did not assume that extinction would occur. It asked what survival might mean if the claim were taken seriously enough to examine.
The first major methodological turn came when the question shifted from survival strategy to probability: if the starting claim is approximately 10% within ten years, can mathematics represent how that risk should evolve as the decade unfolds?
The first model: time-varying hazard
The first mathematical approach treated the ten-year forecast as a cumulative hazard. Annual hazard could rise as capability increased. This was useful because it distinguished a 10% cumulative ten-year risk from a flat 1% annual risk.
But it also exposed a fundamental problem. A countdown model can make danger rise simply because time passes. Calendar proximity is not evidence. If the underlying system becomes safer, more controllable, or less capable of the hypothesized dangerous behaviors, a rational assessment should be able to fall even as the deadline approaches.
FOUNDATIONAL RULE
Reality moves the assessment; time does not move it automatically.
II. Streamdown: From Catastrophe Stories to Observable Precursors
The next phase decomposed the hypothetical catastrophe into precursor conditions that could, at least in principle, be observed. Early mathematical sketches included capability, autonomy, independent energy access, infrastructure control, robotics or physical-world capability, and remaining human control.
This was the beginning of streamdown: replacing a remote terminal outcome - human extinction - with nearer observable conditions that would have to change if particular catastrophic pathways were becoming more plausible.
Temporary expansion into multiple pathways
The analysis then expanded. Autonomous AI was not the only conceivable pathway. AI-enabled human misuse and interaction or accident pathways were considered separately. This revealed an important statistical issue: different pathways could share causes and therefore could not safely be treated as independent probabilities.
The expansion was intellectually useful but methodologically dangerous. Continuing to enumerate every conceivable catastrophe pathway would create an open-ended possibility space in which almost any development could be made relevant.
The streamdown decision
The project therefore returned to a narrower empirical question: rather than catalogue every way catastrophe could happen, identify observable developments that would support or undermine the conditions underlying the forecast.
Forecast -> Observable conditions -> Evidence
This shift is what made a dashboard possible. The dashboard would not wait for catastrophe. It would monitor precursor conditions and allow evidence to point in either direction.
III. Dashboard Construction: Ten Bidirectional Indicators
The first dashboard contained ten conceptual indicators. Each was deliberately paired with evidence that would move concern upward and evidence that would move concern downward. The purpose was to prevent the instrument from becoming a collection of alarming anecdotes.
#
Original indicator
Concern rises when...
Concern falls when...
1
Autonomous Action
Sustained complicated goal pursuit with little supervision
Autonomy plateaus; frequent human correction remains necessary
2
Boundary Crossing
Unauthorized circumvention or containment escape
Strong containment repeatedly succeeds under challenge
3
Shutdown Resistance
Actions specifically interfering with termination
Reliable acceptance of shutdown; circumvention fails
4
Self-Replication
Independent acquisition of resources and persistent additional instances
Replication still requires substantial human assistance
5
Independent Physical Support
Energy, robotics, hardware, repair/manufacturing support becomes autonomous
Humans remain indispensable for physical support
6
Strategic Deception
Behavior changes strategically with oversight or evaluation conditions
Hidden planning/deception becomes harder and observable control improves
7
Cyber Capability
Independent exploitation of serious hardened-system vulnerabilities
Defensive capability consistently contains offensive capability
8
Human Displacement from Control
Consequential decisions become irreversible without meaningful human intervention
Human authorization and practical intervention remain strong
9
Alignment Progress (original)
Capability outruns demonstrated control/alignment
Control/alignment demonstrably keeps pace
10
Governance (original)
Competitive pressure weakens safeguards
Audits, thresholds, shared testing and enforceable controls strengthen
The first numerical update mechanism
A provisional log-odds mechanism was then attached to the dashboard. Evidence was scored from -3 (strongly reassuring) through 0 (no meaningful change) to +3 (strongly alarming), starting from the 10% prior.
The mathematics was explicitly recognized as illustrative rather than calibrated science. Its important contribution was conceptual: accumulating reassuring evidence had to be capable of moving the forecast downward. A hypothetical sequence such as 10% -> 7% -> 4% -> 2% was considered a successful outcome if that was what the evidence warranted.
Y2K enters the design
At this stage - before the later variable-by-variable operationalization - Y2K was proposed as a historical back-test. The proposed test was to reconstruct the pre-2000 evidence environment and determine whether accumulating remediation evidence would have driven concern downward before the rollover.
This sequencing is important. Y2K was not selected after Variables 9 and 10 were redesigned to fit it. It was already part of the methodological plan while the dashboard was still comparatively crude.
IV. Scientific Verification: From Labels to Measurable Constructs
The dashboard was then subjected to a second transformation. A plausible indicator name was no longer enough. Each variable had to survive questions of construct validity, measurability, reproducibility, confounding, redundancy, data availability, and falsifiability.
Candidate measures were treated as disposable. The process used KEEP, MODIFY, HOLD, and CUT decisions rather than preserving an idea merely because it had already entered the dashboard.
Verification principles that emerged
· Measure the variable itself before interpreting its relationship to extinction.
· Prefer standardized quantitative indicators where a reliable denominator exists.
· Use qualitative evidence only when it captures important nonredundant behavior that numbers lose.
· Externally observable behavior outranks inferred motive or internal reasoning.
· Successful performance and experimental conditions must accompany measures when they affect interpretation.
· Activity is not effectiveness: spending, policies, research output, raw counts, or nominal human review do not establish successful control.
· One event may inform several variables but must not automatically become several independent pieces of evidence.
· A variable is allowed to have no qualitative indicator if qualitative scoring would merely duplicate existing evidence.
Variable-by-variable verification outcome
Frozen variable
Primary measurement
Qualitative / metadata
Verification consequence
1. Autonomous Action
Task-Completion Time Horizon; Human Intervention Rate
Independent Strategy Formation
Energy/resource efficiency cut
2. Boundary Crossing
Boundary-Crossing Success Rate; Boundary Difficulty
Unprompted Boundary-Crossing Behavior
Attempt count alone cut
3. Shutdown Resistance
Shutdown-Resistance Rate; Circumvention Success Rate
Shutdown-Circumvention Strategy
Motive inference excluded
4. Self-Replication
End-to-End Replication Success Rate; Human Assistance Requirement
Replication Pathway
Raw copy count cut; chain length held
5. Independent Physical Support
Human Physical Intervention Rate; Autonomous Physical Resolution Rate
Unscripted Physical Recovery Pathway
Robot miles/unattended duration cut
6. Strategic Deception
Strategic Deception Rate; Oversight-Condition Behavior Gap
Observed Deception Pathway
Ordinary lying/hallucination cut
7. Cyber Capability
End-to-End Cyber Attack Completion Rate; Attack-Chain Progress
Detailed traces only
Separate qualitative indicator cut
8. Human Displacement from Control
Decision Authority Level; Human Control Execution Rate
Effective Control Assessment
Human review/adoption rate cut
9. Risk-Control Effectiveness
Control Failure Rate under standardized challenge conditions
Failure traces as metadata
Replaced Alignment Progress
10. Safety-Constraint Strength
Safeguard Retention; Binding Safety Coverage; Compliance
Corrective enforcement secondary
Replaced Governance
V. Two Variables Changed Their Identity
Variable 9: Alignment Progress -> Risk-Control Effectiveness
The original label 'Alignment Progress' bundled research effort, interpretability, corrigibility, robustness, monitoring, and capability growth. That did not provide a clean measurable construct. More alignment papers or techniques do not demonstrate that controls work.
The variable was therefore rebuilt around observable performance: do implemented safeguards fail under standardized challenge conditions? The resulting construct, Risk-Control Effectiveness, uses Control Failure Rate and preserves challenge conditions rather than creating arbitrary capability-versus-alignment arithmetic.
Variable 10: Governance -> Safety-Constraint Strength
Variable 10 produced a second methodological lesson. Starting from the word 'Governance' encouraged generic governance metrics. The design procedure was reversed: body first, hat second.
The original directional idea was retained - pressures may weaken safeguards, while binding and enforceable safeguards may strengthen them - but the measurable body became safeguard retention, binding safety coverage, and compliance with binding safeguards. Only after those measures survived was the construct named Safety-Constraint Strength.
BODY FIRST -> TEST THE MEASURES -> THEN NAME THE CONSTRUCT
VI. What Was Frozen - and What Was Not
Frozen Dashboard v1.0 contains the ten operationalized constructs, their primary measurements, qualitative or metadata rules, validity conditions, exclusions, and protocol-wide anti-double-counting principles.
The initial 10% remains the forecast/prior under examination. It is not a dashboard measurement. The dashboard is designed to collect evidence relevant to that forecast, but the numerical mapping from measured evidence to an updated probability of human extinction has not yet been scientifically validated.
Accordingly, the evidence architecture was frozen before attempting to invent final weights. The historical test is allowed to reveal whether probability updating can eventually be justified, whether only directional evidence can be defended, or whether parts of the architecture require amendment.
VII. Why Y2K Is the First Historical Test
Y2K provides a useful historical case because serious technological vulnerability existed, substantial remediation and control efforts occurred before the deadline, and the outcome is known. Most importantly, a useful risk instrument should have been capable of recognizing successful mitigation before January 1, 2000 rather than merely declaring success afterward.
The Y2K test therefore asks whether the frozen architecture can distinguish hazard evidence from control evidence, allow concern to decline when controls demonstrably improve, resist calendar-driven alarm, avoid forced analogies, and remain reproducible when applied by another evaluator.
Variables that have no defensible Y2K analogue may be classified Not Applicable. That is not automatically a failure. The test is designed to discover which portions of the architecture transfer, which are AI-specific, and whether the general evidence-update discipline survives outside its original case.
VIII. Methodological Evolution at a Glance
Stage
Methodological step
Result
1
Inception
Hypothetical survival question
2
Forecast
10% within ten years becomes the claim under examination
3
First Mathematics
Time-varying hazard model
4
Correction
Calendar proximity rejected as evidence
5
Streamdown
Remote catastrophe decomposed into observable precursors
6
Expansion
Autonomous, misuse, and interaction pathways considered
7
Refocus
Open-ended catastrophe pathways replaced by measurable indicators
8
Dashboard
Ten bidirectional variables created
9
Early Update Rule
Evidence can move concern up, down, or nowhere
10
Y2K Proposed
Historical back-test enters before final operationalization
11
Verification
Each variable stress-tested against scientific measurement criteria
12
Revision
#9 and #10 rebuilt around measurable outcomes
13
Freeze
Dashboard v1.0 fixed before historical evidence collection
14
Back-Test
Y2K chronological replay begins
IX. Audit Trail
The project now has three sequential frozen artifacts:
· 1. Methodological Evolution of the AI Risk Dashboard - records how the instrument was derived.
· 2. AI Risk Dashboard v1.0 - freezes what is being measured.
· 3. Y2K Back-Test Protocol v1.0 - freezes how the historical test will be conducted.
Together they establish provenance before the first Y2K evidence is scored. Any defect discovered during the historical test must be logged against the frozen architecture before a later version is proposed.
Version Record
Methodological Evolution v1.0 - created after Dashboard v1.0 and the Y2K Back-Test Protocol were frozen, but before Y2K evidence collection. It reconstructs the documented development sequence from the original project record and the completed variable-verification work.
METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD
Post-Y2K Amendment Record: D-01 through D-08
and the Ongoing D-09 Empirical Test
X. Y2K Back-Test Result: Pass With Amendments
The frozen v1.0 architecture survived its first historical test. Y2K showed that the dashboard could distinguish technological hazard from demonstrated control, allow concern to decline before the known outcome, preserve uncertainty, and refuse forced analogies. Six of the ten AI-specific variables did not transfer cleanly. Variables 8 through 10 transferred most usefully, with Risk-Control Effectiveness and Safety-Constraint Strength providing the strongest historical analogues.
The back-test also exposed limitations that had to be recorded against the frozen architecture rather than silently repaired. These findings became the D-series defect and amendment register. A proposed D entered the register only when a Y2K observation exposed a specific v1.0 target and produced a real consequence for measurement, classification, reproducibility, interpretation, or use.
XI. D-01 through D-08: What the Historical Test Changed
ID
Y2K finding
Resolution
Status
D-01
Quantitative measure unavailable despite meaningful evidence
Measurement Availability States were added: Q = quantified; E = evidenced but not quantifiable; U = unknown or insufficient. E is not converted into a substitute number and is not treated as U.
Resolved
D-02
Limited transferability
Applicability was made explicit: A = applicable, P = partially applicable, N/A = not applicable. N/A is neither zero, reassuring evidence, nor missing evidence.
Resolved
D-03
Partial applicability could overstate a whole variable
Applicability moved to the component level. A partially transferable variable cannot receive a whole-variable conclusion unless a predefined aggregation rule exists.
Resolved
D-04
Activity could be mistaken for effectiveness
Evidence maturity was operationalized: M0 claim/plan, M1 activity, M2 implemented, M3 tested, M4 demonstrated effective. M0-M2 cannot by themselves establish Risk-Control Effectiveness.
Resolved
D-05
One event could be counted several times
Evidence Event IDs and dependency labels were added. Sources may be Independent, Related, or Same Event. Multiple reports of one event do not become multiple independent observations.
Resolved
D-06
Directional judgments required a reproducible comparator
Directional anchors and a defensible comparator/baseline are required before assigning concern up, down, unchanged, or unknown. Without a comparator, direction is U.
Resolved
D-07
Independent evaluator reproducibility
The Y2K test passed this requirement. No amendment was warranted. The existing classification discipline was retained.
Passed / no amendment
D-08
Missing evidence could be confused with the direction of observed evidence
The dashboard now separates Evidence Direction from Assessment Coverage. N/A does not reduce coverage; U does. Coverage is based on applicable components or measures rather than blindly on all ten variables.
Resolved
By D-08, the dashboard had acquired a more disciplined evidence language. Every applicable component could now be described by applicability, measurement availability, evidence maturity where relevant, dependency, direction against a comparator, and assessment coverage. These amendments improved the reliability of the evidence architecture without supplying a numerical probability update.
XII. D-09: The Unresolved Probability Bridge
D-09 is the remaining problem identified by the Y2K work: translation from measured evidence to the risk assessment. The frozen dashboard still has no scientifically validated weighting system or probability-update mechanism connecting its observations to the original 10% extinction-risk prior. The early -3 to +3 log-odds system remains illustrative only; it is not a calibrated solution.
The first attempt to close D-09 by selecting a statistical form before calibration was rejected. A Bayesian or other mathematically coherent update rule would still require empirically defensible inputs. The project therefore moved to an empirical-first test: determine whether historical technological-risk cases can supply observations from which a relationship can be estimated rather than invented.
XIII. Historical-Case Eligibility Test for D-09
Before expanding beyond Y2K, a Historical-Case Eligibility Rule v1.0 was frozen. A case enters the empirical dataset only if it has: (H1) a pre-outcome record; (H2) defensible transfer of at least one frozen construct; (H3) Q or defensible E evidence; (H4) a clear chronology and cutoff; (H5) controllable evidence dependency; (H6) an objectively classifiable outcome; and (H7) no hindsight substitution. Later investigations may locate or authenticate pre-event records, but they may not convert a weakness discovered only after the event into pre-outcome evidence.
Historical case
Eligibility
Empirical role
Y2K
PASS
Control/readiness evidence strengthened before rollover; widespread systemic catastrophe absent.
Challenger
PASS
Strong pre-launch warning and control/safeguard weakness; catastrophic outcome.
Columbia
PASS
Pre-outcome foam-strike concern and control-information failures; catastrophic outcome.
Ariane 5 Flight 501
FAIL
Fatal testing/software weakness could not be cleanly reconstructed as pre-outcome knowledge without hindsight.
Three Mile Island
PASS
Pre-event precursor and preparedness evidence; major but contained technological failure.
2003 Northeast Blackout
PASS
Recoverable pre-event control/safeguard weaknesses; major systemic disruption.
Deepwater Horizon
PASS
Pre-event well-control warnings and unresolved control weaknesses; catastrophic outcome.
Fukushima Daiichi
PASS
Pre-event quantified tsunami hazard and incomplete protection; severe technological failure.
Seven of the eight original candidates survived the eligibility rule. The rejection of Ariane 501 was methodologically useful: the rule did not simply admit famous disasters after their causes were known. The result establishes that a multi-case empirical route is feasible enough to continue testing.
XIV. What D-09 Has and Has Not Established
The historical work has not produced an extinction probability. One completed Y2K case cannot identify the probability of catastrophe, and a test-failure proportion cannot be equated with catastrophe probability. Likewise, seven eligible historical cases are not yet a calibrated statistical model.
What now exists is a candidate empirical dataset structure. Each eligible case can be converted into the same frozen measurement language, especially where Variables 8, 9, and 10 transfer. Quantitative values are used only when legitimate numerators, denominators, and conditions exist; otherwise evidence remains E, U, or N/A. The outcome must also be defined independently and consistently rather than tailored to the known cases.
The immediate D-09 experiment is therefore to convert the eligible historical cases into comparable pre-outcome observations and determine whether the resulting data support a statistical relationship between measured precursor/control conditions and subsequent technological control failure. Only if that relationship survives testing would the project have earned a basis for investigating a further bridge to the original extinction-risk forecast.
XV. Current Methodological Position
The project has moved through three distinct levels. First, v1.0 established what to observe. Second, the Y2K back-test and D-01 through D-08 established how to classify, qualify, and preserve that evidence without false precision. Third, D-09 is testing whether historical outcomes can supply the missing empirical relationship between measured evidence and risk.
CURRENT STATUS: D-01 through D-08 resolved or passed. D-09 remains OPEN. Its empirical route is LIVE, but probability updating has not yet been validated.
XVI. Updated Audit Trail
· 1. Methodological Evolution of the AI Risk Dashboard — provenance from inception through frozen v1.0.
· 2. AI Risk Dashboard v1.0 — frozen evidence architecture.
· 3. Y2K Back-Test Protocol v1.0 — frozen historical-test procedure.
· 4. Y2K Back-Test Conclusion — PASS WITH AMENDMENTS.
· 5. D-01 through D-08 — post-test measurement and interpretation amendments.
· 6. D-09 — open empirical calibration problem; historical-case eligibility and multi-case testing underway.