Methodological Evaluation AI Risk Dashboard v.1

METHODOLOGICAL EVOLUTION
OF THE AI RISK DASHBOARD

Inception to Frozen v1.0

Pre-Y2K Back-Test Provenance Record

 

Purpose of This Record

This document preserves how the project evolved before the Y2K historical back-test. Its purpose is methodological provenance: to show which ideas came first, which assumptions were challenged, which branches were abandoned, how the ten-variable dashboard emerged, and how its variables were operationalized and stress-tested before historical validation.

This matters because a retrospective test is vulnerable to hindsight. If the instrument were redesigned after examining Y2K, a successful fit could be manufactured. The sequence recorded here establishes that Y2K was proposed as a test of the method before the final variable architecture was completed.

I. Inception: A Survival Thought Experiment Becomes a Forecasting Problem

The project began with a hypothetical survival question prompted by a contemporary claim that advanced AI might present a substantial risk of human extinction within roughly a decade. The initial exercise did not assume that extinction would occur. It asked what survival might mean if the claim were taken seriously enough to examine.

The first major methodological turn came when the question shifted from survival strategy to probability: if the starting claim is approximately 10% within ten years, can mathematics represent how that risk should evolve as the decade unfolds?

The first model: time-varying hazard

The first mathematical approach treated the ten-year forecast as a cumulative hazard. Annual hazard could rise as capability increased. This was useful because it distinguished a 10% cumulative ten-year risk from a flat 1% annual risk.

But it also exposed a fundamental problem. A countdown model can make danger rise simply because time passes. Calendar proximity is not evidence. If the underlying system becomes safer, more controllable, or less capable of the hypothesized dangerous behaviors, a rational assessment should be able to fall even as the deadline approaches.

FOUNDATIONAL RULE

Reality moves the assessment; time does not move it automatically.

II. Streamdown: From Catastrophe Stories to Observable Precursors

The next phase decomposed the hypothetical catastrophe into precursor conditions that could, at least in principle, be observed. Early mathematical sketches included capability, autonomy, independent energy access, infrastructure control, robotics or physical-world capability, and remaining human control.

This was the beginning of streamdown: replacing a remote terminal outcome - human extinction - with nearer observable conditions that would have to change if particular catastrophic pathways were becoming more plausible.

Temporary expansion into multiple pathways

The analysis then expanded. Autonomous AI was not the only conceivable pathway. AI-enabled human misuse and interaction or accident pathways were considered separately. This revealed an important statistical issue: different pathways could share causes and therefore could not safely be treated as independent probabilities.

The expansion was intellectually useful but methodologically dangerous. Continuing to enumerate every conceivable catastrophe pathway would create an open-ended possibility space in which almost any development could be made relevant.

The streamdown decision

The project therefore returned to a narrower empirical question: rather than catalogue every way catastrophe could happen, identify observable developments that would support or undermine the conditions underlying the forecast.

Forecast -> Observable conditions -> Evidence

This shift is what made a dashboard possible. The dashboard would not wait for catastrophe. It would monitor precursor conditions and allow evidence to point in either direction.

III. Dashboard Construction: Ten Bidirectional Indicators

The first dashboard contained ten conceptual indicators. Each was deliberately paired with evidence that would move concern upward and evidence that would move concern downward. The purpose was to prevent the instrument from becoming a collection of alarming anecdotes.

#

Original indicator

Concern rises when...

Concern falls when...

1

Autonomous Action

Sustained complicated goal pursuit with little supervision

Autonomy plateaus; frequent human correction remains necessary

2

Boundary Crossing

Unauthorized circumvention or containment escape

Strong containment repeatedly succeeds under challenge

3

Shutdown Resistance

Actions specifically interfering with termination

Reliable acceptance of shutdown; circumvention fails

4

Self-Replication

Independent acquisition of resources and persistent additional instances

Replication still requires substantial human assistance

5

Independent Physical Support

Energy, robotics, hardware, repair/manufacturing support becomes autonomous

Humans remain indispensable for physical support

6

Strategic Deception

Behavior changes strategically with oversight or evaluation conditions

Hidden planning/deception becomes harder and observable control improves

7

Cyber Capability

Independent exploitation of serious hardened-system vulnerabilities

Defensive capability consistently contains offensive capability

8

Human Displacement from Control

Consequential decisions become irreversible without meaningful human intervention

Human authorization and practical intervention remain strong

9

Alignment Progress (original)

Capability outruns demonstrated control/alignment

Control/alignment demonstrably keeps pace

10

Governance (original)

Competitive pressure weakens safeguards

Audits, thresholds, shared testing and enforceable controls strengthen

The first numerical update mechanism

A provisional log-odds mechanism was then attached to the dashboard. Evidence was scored from -3 (strongly reassuring) through 0 (no meaningful change) to +3 (strongly alarming), starting from the 10% prior.

The mathematics was explicitly recognized as illustrative rather than calibrated science. Its important contribution was conceptual: accumulating reassuring evidence had to be capable of moving the forecast downward. A hypothetical sequence such as 10% -> 7% -> 4% -> 2% was considered a successful outcome if that was what the evidence warranted.

Y2K enters the design

At this stage - before the later variable-by-variable operationalization - Y2K was proposed as a historical back-test. The proposed test was to reconstruct the pre-2000 evidence environment and determine whether accumulating remediation evidence would have driven concern downward before the rollover.

This sequencing is important. Y2K was not selected after Variables 9 and 10 were redesigned to fit it. It was already part of the methodological plan while the dashboard was still comparatively crude.

IV. Scientific Verification: From Labels to Measurable Constructs

The dashboard was then subjected to a second transformation. A plausible indicator name was no longer enough. Each variable had to survive questions of construct validity, measurability, reproducibility, confounding, redundancy, data availability, and falsifiability.

Candidate measures were treated as disposable. The process used KEEP, MODIFY, HOLD, and CUT decisions rather than preserving an idea merely because it had already entered the dashboard.

Verification principles that emerged

·         Measure the variable itself before interpreting its relationship to extinction.

·         Prefer standardized quantitative indicators where a reliable denominator exists.

·         Use qualitative evidence only when it captures important nonredundant behavior that numbers lose.

·         Externally observable behavior outranks inferred motive or internal reasoning.

·         Successful performance and experimental conditions must accompany measures when they affect interpretation.

·         Activity is not effectiveness: spending, policies, research output, raw counts, or nominal human review do not establish successful control.

·         One event may inform several variables but must not automatically become several independent pieces of evidence.

·         A variable is allowed to have no qualitative indicator if qualitative scoring would merely duplicate existing evidence.

Variable-by-variable verification outcome

Frozen variable

Primary measurement

Qualitative / metadata

Verification consequence

1. Autonomous Action

Task-Completion Time Horizon; Human Intervention Rate

Independent Strategy Formation

Energy/resource efficiency cut

2. Boundary Crossing

Boundary-Crossing Success Rate; Boundary Difficulty

Unprompted Boundary-Crossing Behavior

Attempt count alone cut

3. Shutdown Resistance

Shutdown-Resistance Rate; Circumvention Success Rate

Shutdown-Circumvention Strategy

Motive inference excluded

4. Self-Replication

End-to-End Replication Success Rate; Human Assistance Requirement

Replication Pathway

Raw copy count cut; chain length held

5. Independent Physical Support

Human Physical Intervention Rate; Autonomous Physical Resolution Rate

Unscripted Physical Recovery Pathway

Robot miles/unattended duration cut

6. Strategic Deception

Strategic Deception Rate; Oversight-Condition Behavior Gap

Observed Deception Pathway

Ordinary lying/hallucination cut

7. Cyber Capability

End-to-End Cyber Attack Completion Rate; Attack-Chain Progress

Detailed traces only

Separate qualitative indicator cut

8. Human Displacement from Control

Decision Authority Level; Human Control Execution Rate

Effective Control Assessment

Human review/adoption rate cut

9. Risk-Control Effectiveness

Control Failure Rate under standardized challenge conditions

Failure traces as metadata

Replaced Alignment Progress

10. Safety-Constraint Strength

Safeguard Retention; Binding Safety Coverage; Compliance

Corrective enforcement secondary

Replaced Governance

V. Two Variables Changed Their Identity

Variable 9: Alignment Progress -> Risk-Control Effectiveness

The original label 'Alignment Progress' bundled research effort, interpretability, corrigibility, robustness, monitoring, and capability growth. That did not provide a clean measurable construct. More alignment papers or techniques do not demonstrate that controls work.

The variable was therefore rebuilt around observable performance: do implemented safeguards fail under standardized challenge conditions? The resulting construct, Risk-Control Effectiveness, uses Control Failure Rate and preserves challenge conditions rather than creating arbitrary capability-versus-alignment arithmetic.

Variable 10: Governance -> Safety-Constraint Strength

Variable 10 produced a second methodological lesson. Starting from the word 'Governance' encouraged generic governance metrics. The design procedure was reversed: body first, hat second.

The original directional idea was retained - pressures may weaken safeguards, while binding and enforceable safeguards may strengthen them - but the measurable body became safeguard retention, binding safety coverage, and compliance with binding safeguards. Only after those measures survived was the construct named Safety-Constraint Strength.

BODY FIRST -> TEST THE MEASURES -> THEN NAME THE CONSTRUCT

VI. What Was Frozen - and What Was Not

Frozen Dashboard v1.0 contains the ten operationalized constructs, their primary measurements, qualitative or metadata rules, validity conditions, exclusions, and protocol-wide anti-double-counting principles.

The initial 10% remains the forecast/prior under examination. It is not a dashboard measurement. The dashboard is designed to collect evidence relevant to that forecast, but the numerical mapping from measured evidence to an updated probability of human extinction has not yet been scientifically validated.

Accordingly, the evidence architecture was frozen before attempting to invent final weights. The historical test is allowed to reveal whether probability updating can eventually be justified, whether only directional evidence can be defended, or whether parts of the architecture require amendment.

VII. Why Y2K Is the First Historical Test

Y2K provides a useful historical case because serious technological vulnerability existed, substantial remediation and control efforts occurred before the deadline, and the outcome is known. Most importantly, a useful risk instrument should have been capable of recognizing successful mitigation before January 1, 2000 rather than merely declaring success afterward.

The Y2K test therefore asks whether the frozen architecture can distinguish hazard evidence from control evidence, allow concern to decline when controls demonstrably improve, resist calendar-driven alarm, avoid forced analogies, and remain reproducible when applied by another evaluator.

Variables that have no defensible Y2K analogue may be classified Not Applicable. That is not automatically a failure. The test is designed to discover which portions of the architecture transfer, which are AI-specific, and whether the general evidence-update discipline survives outside its original case.

VIII. Methodological Evolution at a Glance

Stage

Methodological step

Result

1

Inception

Hypothetical survival question

2

Forecast

10% within ten years becomes the claim under examination

3

First Mathematics

Time-varying hazard model

4

Correction

Calendar proximity rejected as evidence

5

Streamdown

Remote catastrophe decomposed into observable precursors

6

Expansion

Autonomous, misuse, and interaction pathways considered

7

Refocus

Open-ended catastrophe pathways replaced by measurable indicators

8

Dashboard

Ten bidirectional variables created

9

Early Update Rule

Evidence can move concern up, down, or nowhere

10

Y2K Proposed

Historical back-test enters before final operationalization

11

Verification

Each variable stress-tested against scientific measurement criteria

12

Revision

#9 and #10 rebuilt around measurable outcomes

13

Freeze

Dashboard v1.0 fixed before historical evidence collection

14

Back-Test

Y2K chronological replay begins

IX. Audit Trail

The project now has three sequential frozen artifacts:

·         1. Methodological Evolution of the AI Risk Dashboard - records how the instrument was derived.

·         2. AI Risk Dashboard v1.0 - freezes what is being measured.

·         3. Y2K Back-Test Protocol v1.0 - freezes how the historical test will be conducted.

Together they establish provenance before the first Y2K evidence is scored. Any defect discovered during the historical test must be logged against the frozen architecture before a later version is proposed.

Version Record

Methodological Evolution v1.0 - created after Dashboard v1.0 and the Y2K Back-Test Protocol were frozen, but before Y2K evidence collection. It reconstructs the documented development sequence from the original project record and the completed variable-verification work.

Previous
Previous

Y2K Back-Test Conclusion Statement