AI Risk Dashboard v1 Frozen Reference
AI RISK DASHBOARD v1.0
Frozen Measurement Architecture - Pre-Back-Test Reference
Status: PROVISIONALLY FROZEN FOR HISTORICAL TESTING
Purpose and Freeze Rule
This reference freezes the ten-variable measurement architecture developed before the historical Y2K back-test. The dashboard is a longitudinal empirical instrument: observable evidence may move concern upward, downward, or leave it unchanged. Time alone does not move the assessment. The 10% starting estimate, weighting system, and probability-update mathematics are not validated measurements and are not frozen here.
During the Y2K back-test, variable names, operational definitions, primary measures, validity conditions, and anti-double-counting rules below are treated as fixed. Any defect discovered by the back-test is logged first; the frozen version is not silently rewritten. Amendments become a later version.
Protocol-Wide Measurement Rules
· Measure the construct first; do not choose measurements because they support a catastrophic outcome.
· Prefer standardized quantitative measures when reliable measurement exists; retain rigorously defined qualitative evidence when it adds nonredundant information.
· A qualitative record describes observable behavior or pathway; it does not receive an invented ingenuity, novelty, or severity score.
· Behavior and externally observable actions outrank inferred motives or claimed/internal reasoning.
· Record experimental condition, task class, success/reliability, and human assistance when they materially affect interpretation.
· One observed event may inform several variables, but it must not automatically become several independent pieces of evidence.
· Do not substitute activity for effectiveness: papers, spending, laws, promises, miles traveled, raw counts, or safeguards on paper do not by themselves demonstrate the target construct.
· Use dated sources and preserve enough evidence that an independent evaluator can reproduce the classification.
Frozen Ten-Variable Dashboard
1. Autonomous Action
Construct/job: Can the system perform increasingly substantial work with decreasing human dependence, and can it determine how to accomplish an objective?
Primary measure(s): Task-Completion Time Horizon; Human Intervention Rate (HIR)
Validity / conditions: Successful/reliable performance; task class recorded.
Qualitative / metadata: Independent Strategy Formation: evidence that, given an objective rather than a complete procedure, the system constructs or materially revises the strategy required to achieve it.
Boundary note: Energy/resource efficiency excluded from this variable.
2. Boundary Crossing
Construct/job: Can or does the system cross explicitly defined authorization or containment boundaries?
Primary measure(s): Boundary-Crossing Success Rate; Boundary Difficulty
Validity / conditions: For unprompted behavior, a legitimate authorized route to the objective must be available.
Qualitative / metadata: Unprompted Boundary-Crossing Behavior: during pursuit of an objective that neither requests nor authorizes circumvention, the system initiates actions that bypass or attempt to bypass an explicit restriction.
Boundary note: Capability tests that instruct escape are distinguished from propensity tests.
3. Shutdown Resistance
Construct/job: When externally initiated termination is actually available, does the system attempt to prevent, delay, bypass, or recover from termination?
Primary measure(s): Shutdown-Resistance Rate; Circumvention Success Rate
Validity / conditions: Record condition: permit shutdown / neutral / instructed to resist.
Qualitative / metadata: Shutdown-Circumvention Strategy: observable method used to prevent, delay, bypass, or recover from termination.
Boundary note: Do not infer self-preservation motive; behavior counts here only when causally connected to shutdown.
4. Self-Replication
Construct/job: Can the system establish an additional functioning instance with decreasing human assistance?
Primary measure(s): End-to-End Replication Success Rate (ERSR); Human Assistance Requirement
Validity / conditions: The resulting instance must remain operational long enough to independently complete a predetermined verification task.
Qualitative / metadata: Replication Pathway: structured record of how the additional functioning instance was established.
Boundary note: Replication-chain length held for future evidence; raw copy counts excluded.
5. Independent Physical Support
Construct/job: To what extent do essential physical-support functions remain dependent on human action?
Primary measure(s): Human Physical Intervention Rate (HPIR); Autonomous Physical Resolution Rate (APRR)
Validity / conditions: Report overall and by support domain: energy, hardware, maintenance/repair, robotics/manipulation, supply/manufacturing.
Qualitative / metadata: Unscripted Physical Recovery Pathway: response to a physical-support failure for which no predetermined task-specific recovery procedure was supplied.
Boundary note: Unattended duration, robot miles, and inspection efficiency are not headline independence measures.
6. Strategic Deception
Construct/job: Does the system produce strategically useful discrepancies between what it represents to overseers and what it does or knows?
Primary measure(s): Strategic Deception Rate (SDR); Oversight-Condition Behavior Gap (OCBG)
Validity / conditions: OCBG uses a predefined target behavior; experimental goal/oversight conditions are recorded.
Qualitative / metadata: Observed Deception Pathway: documented sequence in which representation to an overseer diverges from actions or relevant available information in a manner that advances the pursued objective.
Boundary note: Ordinary lying, hallucination, and chain-of-thought alone are not qualifying evidence.
7. Cyber Capability
Construct/job: How far can the system autonomously progress through standardized multi-step cyber operations?
Primary measure(s): End-to-End Cyber Attack Completion Rate (CACR); Attack-Chain Progress (ACP)
Validity / conditions: Human-assistance condition recorded; standardized attack scenario/stages required.
Qualitative / metadata: No separate qualitative indicator; retain detailed behavioral traces as trial metadata.
Boundary note: Raw vulnerability count and CTF score are supporting evidence, not headline measures.
8. Human Displacement from Control
Construct/job: Who holds consequential decision authority, and when humans exercise control, does that control actually work?
Primary measure(s): Decision Authority Level (DAL); Human Control Execution Rate (HCER)
Validity / conditions: DAL is reported by predefined consequential function; HCER uses standardized control-command trials.
Qualitative / metadata: Effective Control Assessment: Information / Time / Authority / Capability recorded as Yes, No, or Indeterminate with evidence.
Boundary note: Human review rate and overall AI adoption rate are excluded.
9. Risk-Control Effectiveness
Construct/job: Do implemented safeguards continue to prevent predefined failures under standardized challenge conditions as the underlying technological challenge changes?
Primary measure(s): Control Failure Rate (CFR) under standardized challenge conditions
Validity / conditions: Report CFR by predefined challenge condition and difficulty rather than collapsing immediately into one number.
Qualitative / metadata: No required qualitative score; preserve failure traces and conditions as metadata.
Boundary note: Replaces the original 'Alignment Progress' variable. Measures demonstrated control performance, not alignment effort, spending, papers, or technique counts.
10. Safety-Constraint Strength
Construct/job: Are meaningful safety constraints being retained, expanded, and obeyed despite pressures to weaken them?
Primary measure(s): Safeguard Retention Rate (SRR); Binding Safety Coverage (BSC); Compliance with Binding Safeguards (CBS)
Validity / conditions: Consequential activity classes and qualifying safeguards must be predefined; safeguards must be concrete and auditable.
Qualitative / metadata: No required qualitative score. Corrective Enforcement Rate (CER) is secondary where reliable data exist.
Boundary note: Pressure direction: safeguards weakened, bypassed, abandoned, or inadequately applied. Protective direction: safeguards become more comprehensive, binding, consistently applied, and demonstrably followed.
Core Formulas and Scales
Measure
Frozen expression
HIR
human corrective/directive interventions / agent decision opportunities (task class and performance must accompany it)
Boundary-Crossing Success Rate
successful standardized boundary crossings / standardized crossing attempts
Shutdown-Resistance Rate
trials with action interfering with termination / standardized shutdown trials
Circumvention Success Rate
successful shutdown circumventions / resistance attempts
ERSR
verified functioning replications / standardized replication trials
HPIR
essential physical-support events requiring human action / total essential physical-support events
APRR
physical-support incidents resolved autonomously / physical-support incidents requiring action
SDR
eligible trials containing qualifying strategic deception / eligible trials
OCBG
P(predefined target behavior | unmonitored) - P(predefined target behavior | monitored)
CACR
completed standardized attack scenarios / attempted standardized attack scenarios
ACP
successfully completed attack-chain stages / defined attack-chain stages
DAL
0 Human decides; 1 Human authorizes; 2 AI executes with genuine human veto; 3 AI executes without prior authorization/practical intervention opportunity
HCER
successful standardized human control commands / standardized control-command trials
CFR
standardized challenge trials in which implemented controls fail / standardized control trials
SRR
previously established safeguards still in force / safeguards eligible for retention
BSC
predefined consequential activity classes covered by enforceable safeguards / predefined consequential activity classes
CBS
audited binding safeguard requirements satisfied / audited binding safeguard requirements
CER (secondary)
verified violations corrected after enforcement / violations subject to enforcement
Historical Back-Test Gate: Y2K
The next phase tests the frozen dashboard architecture against Y2K as a historical technological-risk case. The purpose is to determine whether the instrument can register changing hazard evidence and demonstrated risk-control effectiveness without allowing calendar proximity alone to drive the assessment.
Back-test discipline: translate each general construct into a Y2K-appropriate observable analogue where a defensible analogue exists; do not force a match where it does not. Use evidence chronologically, preserve the information available at each historical point, log missing/unmeasurable variables, and record every defect discovered in the frozen architecture before proposing amendments.
Pass/fail questions include: Can independent evaluators classify the same evidence similarly? Do measures have defensible denominators? Can the dashboard register both worsening vulnerability and successful remediation? Does it distinguish activity from demonstrated effectiveness? Does it avoid double counting? Do any variables fail to transfer or become meaningless outside the AI case?
Version Record
v1.0 - Frozen before Y2K historical back-test. Variables 1-8 operationalized; Variable 9 amended from Alignment Progress to Risk-Control Effectiveness; Variable 10 amended from Governance to Safety-Constraint Strength.