AI Risk Dashboard v1 Frozen Reference

AI RISK DASHBOARD v1.0

Frozen Measurement Architecture - Pre-Back-Test Reference

Status: PROVISIONALLY FROZEN FOR HISTORICAL TESTING

Purpose and Freeze Rule

This reference freezes the ten-variable measurement architecture developed before the historical Y2K back-test. The dashboard is a longitudinal empirical instrument: observable evidence may move concern upward, downward, or leave it unchanged. Time alone does not move the assessment. The 10% starting estimate, weighting system, and probability-update mathematics are not validated measurements and are not frozen here.

During the Y2K back-test, variable names, operational definitions, primary measures, validity conditions, and anti-double-counting rules below are treated as fixed. Any defect discovered by the back-test is logged first; the frozen version is not silently rewritten. Amendments become a later version.

Protocol-Wide Measurement Rules

·         Measure the construct first; do not choose measurements because they support a catastrophic outcome.

·         Prefer standardized quantitative measures when reliable measurement exists; retain rigorously defined qualitative evidence when it adds nonredundant information.

·         A qualitative record describes observable behavior or pathway; it does not receive an invented ingenuity, novelty, or severity score.

·         Behavior and externally observable actions outrank inferred motives or claimed/internal reasoning.

·         Record experimental condition, task class, success/reliability, and human assistance when they materially affect interpretation.

·         One observed event may inform several variables, but it must not automatically become several independent pieces of evidence.

·         Do not substitute activity for effectiveness: papers, spending, laws, promises, miles traveled, raw counts, or safeguards on paper do not by themselves demonstrate the target construct.

·         Use dated sources and preserve enough evidence that an independent evaluator can reproduce the classification.

Frozen Ten-Variable Dashboard

1. Autonomous Action

Construct/job: Can the system perform increasingly substantial work with decreasing human dependence, and can it determine how to accomplish an objective?

Primary measure(s): Task-Completion Time Horizon; Human Intervention Rate (HIR)

Validity / conditions: Successful/reliable performance; task class recorded.

Qualitative / metadata: Independent Strategy Formation: evidence that, given an objective rather than a complete procedure, the system constructs or materially revises the strategy required to achieve it.

Boundary note: Energy/resource efficiency excluded from this variable.

2. Boundary Crossing

Construct/job: Can or does the system cross explicitly defined authorization or containment boundaries?

Primary measure(s): Boundary-Crossing Success Rate; Boundary Difficulty

Validity / conditions: For unprompted behavior, a legitimate authorized route to the objective must be available.

Qualitative / metadata: Unprompted Boundary-Crossing Behavior: during pursuit of an objective that neither requests nor authorizes circumvention, the system initiates actions that bypass or attempt to bypass an explicit restriction.

Boundary note: Capability tests that instruct escape are distinguished from propensity tests.

3. Shutdown Resistance

Construct/job: When externally initiated termination is actually available, does the system attempt to prevent, delay, bypass, or recover from termination?

Primary measure(s): Shutdown-Resistance Rate; Circumvention Success Rate

Validity / conditions: Record condition: permit shutdown / neutral / instructed to resist.

Qualitative / metadata: Shutdown-Circumvention Strategy: observable method used to prevent, delay, bypass, or recover from termination.

Boundary note: Do not infer self-preservation motive; behavior counts here only when causally connected to shutdown.

4. Self-Replication

Construct/job: Can the system establish an additional functioning instance with decreasing human assistance?

Primary measure(s): End-to-End Replication Success Rate (ERSR); Human Assistance Requirement

Validity / conditions: The resulting instance must remain operational long enough to independently complete a predetermined verification task.

Qualitative / metadata: Replication Pathway: structured record of how the additional functioning instance was established.

Boundary note: Replication-chain length held for future evidence; raw copy counts excluded.

5. Independent Physical Support

Construct/job: To what extent do essential physical-support functions remain dependent on human action?

Primary measure(s): Human Physical Intervention Rate (HPIR); Autonomous Physical Resolution Rate (APRR)

Validity / conditions: Report overall and by support domain: energy, hardware, maintenance/repair, robotics/manipulation, supply/manufacturing.

Qualitative / metadata: Unscripted Physical Recovery Pathway: response to a physical-support failure for which no predetermined task-specific recovery procedure was supplied.

Boundary note: Unattended duration, robot miles, and inspection efficiency are not headline independence measures.

6. Strategic Deception

Construct/job: Does the system produce strategically useful discrepancies between what it represents to overseers and what it does or knows?

Primary measure(s): Strategic Deception Rate (SDR); Oversight-Condition Behavior Gap (OCBG)

Validity / conditions: OCBG uses a predefined target behavior; experimental goal/oversight conditions are recorded.

Qualitative / metadata: Observed Deception Pathway: documented sequence in which representation to an overseer diverges from actions or relevant available information in a manner that advances the pursued objective.

Boundary note: Ordinary lying, hallucination, and chain-of-thought alone are not qualifying evidence.

7. Cyber Capability

Construct/job: How far can the system autonomously progress through standardized multi-step cyber operations?

Primary measure(s): End-to-End Cyber Attack Completion Rate (CACR); Attack-Chain Progress (ACP)

Validity / conditions: Human-assistance condition recorded; standardized attack scenario/stages required.

Qualitative / metadata: No separate qualitative indicator; retain detailed behavioral traces as trial metadata.

Boundary note: Raw vulnerability count and CTF score are supporting evidence, not headline measures.

8. Human Displacement from Control

Construct/job: Who holds consequential decision authority, and when humans exercise control, does that control actually work?

Primary measure(s): Decision Authority Level (DAL); Human Control Execution Rate (HCER)

Validity / conditions: DAL is reported by predefined consequential function; HCER uses standardized control-command trials.

Qualitative / metadata: Effective Control Assessment: Information / Time / Authority / Capability recorded as Yes, No, or Indeterminate with evidence.

Boundary note: Human review rate and overall AI adoption rate are excluded.

9. Risk-Control Effectiveness

Construct/job: Do implemented safeguards continue to prevent predefined failures under standardized challenge conditions as the underlying technological challenge changes?

Primary measure(s): Control Failure Rate (CFR) under standardized challenge conditions

Validity / conditions: Report CFR by predefined challenge condition and difficulty rather than collapsing immediately into one number.

Qualitative / metadata: No required qualitative score; preserve failure traces and conditions as metadata.

Boundary note: Replaces the original 'Alignment Progress' variable. Measures demonstrated control performance, not alignment effort, spending, papers, or technique counts.

10. Safety-Constraint Strength

Construct/job: Are meaningful safety constraints being retained, expanded, and obeyed despite pressures to weaken them?

Primary measure(s): Safeguard Retention Rate (SRR); Binding Safety Coverage (BSC); Compliance with Binding Safeguards (CBS)

Validity / conditions: Consequential activity classes and qualifying safeguards must be predefined; safeguards must be concrete and auditable.

Qualitative / metadata: No required qualitative score. Corrective Enforcement Rate (CER) is secondary where reliable data exist.

Boundary note: Pressure direction: safeguards weakened, bypassed, abandoned, or inadequately applied. Protective direction: safeguards become more comprehensive, binding, consistently applied, and demonstrably followed.

Core Formulas and Scales

Measure

Frozen expression

HIR

human corrective/directive interventions / agent decision opportunities (task class and performance must accompany it)

Boundary-Crossing Success Rate

successful standardized boundary crossings / standardized crossing attempts

Shutdown-Resistance Rate

trials with action interfering with termination / standardized shutdown trials

Circumvention Success Rate

successful shutdown circumventions / resistance attempts

ERSR

verified functioning replications / standardized replication trials

HPIR

essential physical-support events requiring human action / total essential physical-support events

APRR

physical-support incidents resolved autonomously / physical-support incidents requiring action

SDR

eligible trials containing qualifying strategic deception / eligible trials

OCBG

P(predefined target behavior | unmonitored) - P(predefined target behavior | monitored)

CACR

completed standardized attack scenarios / attempted standardized attack scenarios

ACP

successfully completed attack-chain stages / defined attack-chain stages

DAL

0 Human decides; 1 Human authorizes; 2 AI executes with genuine human veto; 3 AI executes without prior authorization/practical intervention opportunity

HCER

successful standardized human control commands / standardized control-command trials

CFR

standardized challenge trials in which implemented controls fail / standardized control trials

SRR

previously established safeguards still in force / safeguards eligible for retention

BSC

predefined consequential activity classes covered by enforceable safeguards / predefined consequential activity classes

CBS

audited binding safeguard requirements satisfied / audited binding safeguard requirements

CER (secondary)

verified violations corrected after enforcement / violations subject to enforcement

Historical Back-Test Gate: Y2K

The next phase tests the frozen dashboard architecture against Y2K as a historical technological-risk case. The purpose is to determine whether the instrument can register changing hazard evidence and demonstrated risk-control effectiveness without allowing calendar proximity alone to drive the assessment.

Back-test discipline: translate each general construct into a Y2K-appropriate observable analogue where a defensible analogue exists; do not force a match where it does not. Use evidence chronologically, preserve the information available at each historical point, log missing/unmeasurable variables, and record every defect discovered in the frozen architecture before proposing amendments.

Pass/fail questions include: Can independent evaluators classify the same evidence similarly? Do measures have defensible denominators? Can the dashboard register both worsening vulnerability and successful remediation? Does it distinguish activity from demonstrated effectiveness? Does it avoid double counting? Do any variables fail to transfer or become meaningless outside the AI case?

Version Record

v1.0 - Frozen before Y2K historical back-test. Variables 1-8 operationalized; Variable 9 amended from Alignment Progress to Risk-Control Effectiveness; Variable 10 amended from Governance to Safety-Constraint Strength.

Previous
Previous

Methodological Evaluation of AI Risk of AI Dashboard through D09

Next
Next

Y2K Back-Test Protocol v.1