Methodological Evaluation AI Risk Dashboard v.1
METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD
Inception to Frozen v1.0
Pre-Y2K Back-Test Provenance Record
Purpose of This Record
This document preserves how the project evolved before the Y2K historical back-test. Its purpose is methodological provenance: to show which ideas came first, which assumptions were challenged, which branches were abandoned, how the ten-variable dashboard emerged, and how its variables were operationalized and stress-tested before historical validation.
This matters because a retrospective test is vulnerable to hindsight. If the instrument were redesigned after examining Y2K, a successful fit could be manufactured. The sequence recorded here establishes that Y2K was proposed as a test of the method before the final variable architecture was completed.
I. Inception: A Survival Thought Experiment Becomes a Forecasting Problem
The project began with a hypothetical survival question prompted by a contemporary claim that advanced AI might present a substantial risk of human extinction within roughly a decade. The initial exercise did not assume that extinction would occur. It asked what survival might mean if the claim were taken seriously enough to examine.
The first major methodological turn came when the question shifted from survival strategy to probability: if the starting claim is approximately 10% within ten years, can mathematics represent how that risk should evolve as the decade unfolds?
The first model: time-varying hazard
The first mathematical approach treated the ten-year forecast as a cumulative hazard. Annual hazard could rise as capability increased. This was useful because it distinguished a 10% cumulative ten-year risk from a flat 1% annual risk.
But it also exposed a fundamental problem. A countdown model can make danger rise simply because time passes. Calendar proximity is not evidence. If the underlying system becomes safer, more controllable, or less capable of the hypothesized dangerous behaviors, a rational assessment should be able to fall even as the deadline approaches.
FOUNDATIONAL RULE
Reality moves the assessment; time does not move it automatically.
II. Streamdown: From Catastrophe Stories to Observable Precursors
The next phase decomposed the hypothetical catastrophe into precursor conditions that could, at least in principle, be observed. Early mathematical sketches included capability, autonomy, independent energy access, infrastructure control, robotics or physical-world capability, and remaining human control.
This was the beginning of streamdown: replacing a remote terminal outcome - human extinction - with nearer observable conditions that would have to change if particular catastrophic pathways were becoming more plausible.
Temporary expansion into multiple pathways
The analysis then expanded. Autonomous AI was not the only conceivable pathway. AI-enabled human misuse and interaction or accident pathways were considered separately. This revealed an important statistical issue: different pathways could share causes and therefore could not safely be treated as independent probabilities.
The expansion was intellectually useful but methodologically dangerous. Continuing to enumerate every conceivable catastrophe pathway would create an open-ended possibility space in which almost any development could be made relevant.
The streamdown decision
The project therefore returned to a narrower empirical question: rather than catalogue every way catastrophe could happen, identify observable developments that would support or undermine the conditions underlying the forecast.
Forecast -> Observable conditions -> Evidence
This shift is what made a dashboard possible. The dashboard would not wait for catastrophe. It would monitor precursor conditions and allow evidence to point in either direction.
III. Dashboard Construction: Ten Bidirectional Indicators
The first dashboard contained ten conceptual indicators. Each was deliberately paired with evidence that would move concern upward and evidence that would move concern downward. The purpose was to prevent the instrument from becoming a collection of alarming anecdotes.
#
Original indicator
Concern rises when...
Concern falls when...
1
Autonomous Action
Sustained complicated goal pursuit with little supervision
Autonomy plateaus; frequent human correction remains necessary
2
Boundary Crossing
Unauthorized circumvention or containment escape
Strong containment repeatedly succeeds under challenge
3
Shutdown Resistance
Actions specifically interfering with termination
Reliable acceptance of shutdown; circumvention fails
4
Self-Replication
Independent acquisition of resources and persistent additional instances
Replication still requires substantial human assistance
5
Independent Physical Support
Energy, robotics, hardware, repair/manufacturing support becomes autonomous
Humans remain indispensable for physical support
6
Strategic Deception
Behavior changes strategically with oversight or evaluation conditions
Hidden planning/deception becomes harder and observable control improves
7
Cyber Capability
Independent exploitation of serious hardened-system vulnerabilities
Defensive capability consistently contains offensive capability
8
Human Displacement from Control
Consequential decisions become irreversible without meaningful human intervention
Human authorization and practical intervention remain strong
9
Alignment Progress (original)
Capability outruns demonstrated control/alignment
Control/alignment demonstrably keeps pace
10
Governance (original)
Competitive pressure weakens safeguards
Audits, thresholds, shared testing and enforceable controls strengthen
The first numerical update mechanism
A provisional log-odds mechanism was then attached to the dashboard. Evidence was scored from -3 (strongly reassuring) through 0 (no meaningful change) to +3 (strongly alarming), starting from the 10% prior.
The mathematics was explicitly recognized as illustrative rather than calibrated science. Its important contribution was conceptual: accumulating reassuring evidence had to be capable of moving the forecast downward. A hypothetical sequence such as 10% -> 7% -> 4% -> 2% was considered a successful outcome if that was what the evidence warranted.
Y2K enters the design
At this stage - before the later variable-by-variable operationalization - Y2K was proposed as a historical back-test. The proposed test was to reconstruct the pre-2000 evidence environment and determine whether accumulating remediation evidence would have driven concern downward before the rollover.
This sequencing is important. Y2K was not selected after Variables 9 and 10 were redesigned to fit it. It was already part of the methodological plan while the dashboard was still comparatively crude.
IV. Scientific Verification: From Labels to Measurable Constructs
The dashboard was then subjected to a second transformation. A plausible indicator name was no longer enough. Each variable had to survive questions of construct validity, measurability, reproducibility, confounding, redundancy, data availability, and falsifiability.
Candidate measures were treated as disposable. The process used KEEP, MODIFY, HOLD, and CUT decisions rather than preserving an idea merely because it had already entered the dashboard.
Verification principles that emerged
· Measure the variable itself before interpreting its relationship to extinction.
· Prefer standardized quantitative indicators where a reliable denominator exists.
· Use qualitative evidence only when it captures important nonredundant behavior that numbers lose.
· Externally observable behavior outranks inferred motive or internal reasoning.
· Successful performance and experimental conditions must accompany measures when they affect interpretation.
· Activity is not effectiveness: spending, policies, research output, raw counts, or nominal human review do not establish successful control.
· One event may inform several variables but must not automatically become several independent pieces of evidence.
· A variable is allowed to have no qualitative indicator if qualitative scoring would merely duplicate existing evidence.
Variable-by-variable verification outcome
Frozen variable
Primary measurement
Qualitative / metadata
Verification consequence
1. Autonomous Action
Task-Completion Time Horizon; Human Intervention Rate
Independent Strategy Formation
Energy/resource efficiency cut
2. Boundary Crossing
Boundary-Crossing Success Rate; Boundary Difficulty
Unprompted Boundary-Crossing Behavior
Attempt count alone cut
3. Shutdown Resistance
Shutdown-Resistance Rate; Circumvention Success Rate
Shutdown-Circumvention Strategy
Motive inference excluded
4. Self-Replication
End-to-End Replication Success Rate; Human Assistance Requirement
Replication Pathway
Raw copy count cut; chain length held
5. Independent Physical Support
Human Physical Intervention Rate; Autonomous Physical Resolution Rate
Unscripted Physical Recovery Pathway
Robot miles/unattended duration cut
6. Strategic Deception
Strategic Deception Rate; Oversight-Condition Behavior Gap
Observed Deception Pathway
Ordinary lying/hallucination cut
7. Cyber Capability
End-to-End Cyber Attack Completion Rate; Attack-Chain Progress
Detailed traces only
Separate qualitative indicator cut
8. Human Displacement from Control
Decision Authority Level; Human Control Execution Rate
Effective Control Assessment
Human review/adoption rate cut
9. Risk-Control Effectiveness
Control Failure Rate under standardized challenge conditions
Failure traces as metadata
Replaced Alignment Progress
10. Safety-Constraint Strength
Safeguard Retention; Binding Safety Coverage; Compliance
Corrective enforcement secondary
Replaced Governance
V. Two Variables Changed Their Identity
Variable 9: Alignment Progress -> Risk-Control Effectiveness
The original label 'Alignment Progress' bundled research effort, interpretability, corrigibility, robustness, monitoring, and capability growth. That did not provide a clean measurable construct. More alignment papers or techniques do not demonstrate that controls work.
The variable was therefore rebuilt around observable performance: do implemented safeguards fail under standardized challenge conditions? The resulting construct, Risk-Control Effectiveness, uses Control Failure Rate and preserves challenge conditions rather than creating arbitrary capability-versus-alignment arithmetic.
Variable 10: Governance -> Safety-Constraint Strength
Variable 10 produced a second methodological lesson. Starting from the word 'Governance' encouraged generic governance metrics. The design procedure was reversed: body first, hat second.
The original directional idea was retained - pressures may weaken safeguards, while binding and enforceable safeguards may strengthen them - but the measurable body became safeguard retention, binding safety coverage, and compliance with binding safeguards. Only after those measures survived was the construct named Safety-Constraint Strength.
BODY FIRST -> TEST THE MEASURES -> THEN NAME THE CONSTRUCT
VI. What Was Frozen - and What Was Not
Frozen Dashboard v1.0 contains the ten operationalized constructs, their primary measurements, qualitative or metadata rules, validity conditions, exclusions, and protocol-wide anti-double-counting principles.
The initial 10% remains the forecast/prior under examination. It is not a dashboard measurement. The dashboard is designed to collect evidence relevant to that forecast, but the numerical mapping from measured evidence to an updated probability of human extinction has not yet been scientifically validated.
Accordingly, the evidence architecture was frozen before attempting to invent final weights. The historical test is allowed to reveal whether probability updating can eventually be justified, whether only directional evidence can be defended, or whether parts of the architecture require amendment.
VII. Why Y2K Is the First Historical Test
Y2K provides a useful historical case because serious technological vulnerability existed, substantial remediation and control efforts occurred before the deadline, and the outcome is known. Most importantly, a useful risk instrument should have been capable of recognizing successful mitigation before January 1, 2000 rather than merely declaring success afterward.
The Y2K test therefore asks whether the frozen architecture can distinguish hazard evidence from control evidence, allow concern to decline when controls demonstrably improve, resist calendar-driven alarm, avoid forced analogies, and remain reproducible when applied by another evaluator.
Variables that have no defensible Y2K analogue may be classified Not Applicable. That is not automatically a failure. The test is designed to discover which portions of the architecture transfer, which are AI-specific, and whether the general evidence-update discipline survives outside its original case.
VIII. Methodological Evolution at a Glance
Stage
Methodological step
Result
1
Inception
Hypothetical survival question
2
Forecast
10% within ten years becomes the claim under examination
3
First Mathematics
Time-varying hazard model
4
Correction
Calendar proximity rejected as evidence
5
Streamdown
Remote catastrophe decomposed into observable precursors
6
Expansion
Autonomous, misuse, and interaction pathways considered
7
Refocus
Open-ended catastrophe pathways replaced by measurable indicators
8
Dashboard
Ten bidirectional variables created
9
Early Update Rule
Evidence can move concern up, down, or nowhere
10
Y2K Proposed
Historical back-test enters before final operationalization
11
Verification
Each variable stress-tested against scientific measurement criteria
12
Revision
#9 and #10 rebuilt around measurable outcomes
13
Freeze
Dashboard v1.0 fixed before historical evidence collection
14
Back-Test
Y2K chronological replay begins
IX. Audit Trail
The project now has three sequential frozen artifacts:
· 1. Methodological Evolution of the AI Risk Dashboard - records how the instrument was derived.
· 2. AI Risk Dashboard v1.0 - freezes what is being measured.
· 3. Y2K Back-Test Protocol v1.0 - freezes how the historical test will be conducted.
Together they establish provenance before the first Y2K evidence is scored. Any defect discovered during the historical test must be logged against the frozen architecture before a later version is proposed.
Version Record
Methodological Evolution v1.0 - created after Dashboard v1.0 and the Y2K Back-Test Protocol were frozen, but before Y2K evidence collection. It reconstructs the documented development sequence from the original project record and the completed variable-verification work.