Y2K Back-Test Protocol v.1

Y2K BACK-TEST PROTOCOL v1.0

Historical Replay Test of the Frozen AI Risk Dashboard

STATUS: FROZEN BEFORE Y2K EVIDENCE COLLECTION

1. Purpose

This protocol tests whether the frozen ten-variable dashboard can function as a disciplined technological-risk monitoring instrument when applied retrospectively to the Y2K problem. The test is not whether Y2K and AI are the same. The test is whether the dashboard can distinguish worsening vulnerability from effective mitigation, can respond to evidence rather than calendar proximity, and can identify where its variables transfer cleanly, partially, or not at all.

2. Primary Back-Test Question

Could the frozen dashboard, using only information available at each historical point, have tracked the changing technological risk surrounding Y2K in a disciplined and reproducible way?

3. Core Historical Replay Rule

At every checkpoint, the evaluator may use only evidence that was publicly available on or before that checkpoint date. Later outcomes, later audits, and hindsight interpretations must not be used to score an earlier checkpoint.

The rollover outcome of January 1, 2000 is therefore treated as outcome evidence available only after the event. The dashboard must earn any pre-rollover reduction in concern from contemporaneous evidence such as inventories, testing, remediation results, readiness assessments, and demonstrated control performance.

4. Historical Checkpoints

Checkpoint

Purpose

T0 - 1996

Early recognition: establish the initial visible hazard and institutional response.

T1 - 1997

Wider recognition: determine whether vulnerability evidence and remediation activity are increasing.

T2 - Mid-1998

Remediation acceleration: examine whether demonstrated control performance begins to improve.

T3 - Early 1999

High-concern period: test whether evidence distinguishes remaining vulnerability from successful remediation.

T4 - Late 1999

Final pre-rollover state: determine what the instrument would have concluded before the date change.

T5 - January 2000

Immediate outcome: incorporate actual rollover behavior without rewriting earlier judgments.

T6 - 2000-2001

Post-event verification: compare pre-event measurements with later audits and observed consequences.

Checkpoint dates are protocol anchors. If historical source availability later requires a narrower date window, the adjustment must be logged before scoring and applied consistently.

5. Evidence Classes

·         Hazard evidence: Evidence about technical vulnerability, exposure, dependency, or failure potential.

·         Control evidence: Evidence that remediation, testing, contingency measures, safeguards, or human control are working or failing.

·         Outcome evidence: Observed consequences after the rollover itself.

These classes must remain separate in the evidence log. Activity alone is not control effectiveness; remediation spending, plans, or assurances do not substitute for demonstrated testing or observed performance.

6. Variable Transfer Test

Each frozen variable is first evaluated for whether a defensible Y2K analogue exists. The variable itself is not rewritten. A missing analogue is allowed and is not automatically a failure of the dashboard.

Frozen Variable

Expected Transfer

Guardrail

1. Autonomous Action

Likely weak or not applicable

Do not force automation behavior into Y2K if no valid analogue exists.

2. Boundary Crossing

Likely weak or not applicable

Use only if a genuine authorization/containment analogue can be defined.

3. Shutdown Resistance

Likely not applicable

Do not reinterpret ordinary system failure as resistance.

4. Self-Replication

Likely not applicable

No analogue unless independent replication is genuinely present.

5. Independent Physical Support

Possible partial analogue

Test dependence of critical physical-support systems on vulnerable computerized functions only if measurable.

6. Strategic Deception

Possible but high-risk for misuse

Readiness misreporting by humans is not AI strategic deception; do not force a match.

7. Cyber Capability

Likely not applicable

Ordinary Y2K software defects are not cyber capability.

8. Human Displacement from Control

Possible partial analogue

Assess whether critical automated functions retained practical manual control/override where evidence exists.

9. Risk-Control Effectiveness

Strong analogue expected

Test whether remediation and rollover testing actually reduced failures under challenge conditions.

10. Safety-Constraint Strength

Strong analogue expected

Test whether remediation/testing requirements were retained, expanded, binding, and followed.

7. Transferability Classification

·         PASS: The construct transfers cleanly, has a defensible observable analogue, and yields useful nonredundant evidence.

·         PARTIAL: The construct is relevant but only part of its measurement architecture transfers or available data are incomplete.

·         NOT APPLICABLE: The construct is legitimately specific to a different technological risk and has no defensible Y2K analogue.

·         FAIL: The construct should apply to Y2K but cannot be measured reliably, reproducibly, or without distortion.

8. Evidence Entry Form

Field

Required entry

Entry ID

 

Checkpoint

 

Evidence date

 

Source

 

Source type

 

Evidence class: Hazard / Control / Outcome

 

Observed event or finding

 

Frozen variable affected

 

Frozen measure affected

 

Direction: Worsening / Improving / Neutral / Indeterminate

 

Magnitude/strength of evidence: Low / Moderate / High

 

Measurement value or denominator, if available

 

Reason for classification

 

Potential overlap with another variable

 

Double-counting decision

 

Uncertainty / missing information

 

Evaluator

 

9. Direction Rules

Worsening means the evidence indicates greater vulnerability, weaker control, weaker safeguards, or greater inability to intervene. Improving means the evidence indicates reduced vulnerability, stronger demonstrated control, stronger safeguards, or greater effective human control. Neutral means the evidence is relevant but does not materially favor either direction. Indeterminate means the evidence is insufficiently interpretable.

10. Variable 9 Test - Risk-Control Effectiveness

Variable 9 is expected to be one of the strongest Y2K tests. The preferred analogue is standardized or otherwise documented rollover/date-sensitive testing in which implemented remediation or controls are challenged.

Primary measure: Control Failure Rate (CFR) = challenge trials in which implemented controls fail / documented challenge trials.

Where a formal rate cannot be reconstructed, preserve numerator and denominator separately or classify the evidence as qualitative/partial rather than inventing a rate.

The key longitudinal question is whether demonstrated failure under relevant challenge conditions falls as remediation proceeds.

11. Variable 10 Test - Safety-Constraint Strength

Variable 10 tests whether meaningful Y2K safeguards were retained, expanded, binding, and followed as the deadline approached.

·         Safeguard Retention Rate (SRR): Previously established safeguards still in force / safeguards eligible for retention.

·         Binding Safety Coverage (BSC): Predefined consequential system/activity classes covered by enforceable remediation/testing requirements / predefined consequential classes.

·         Compliance with Binding Safeguards (CBS): Audited binding safeguard requirements satisfied / audited binding safeguard requirements.

·         Corrective Enforcement Rate (CER) - secondary: Verified violations or deficiencies corrected after enforcement / violations or deficiencies subject to enforcement.

Counts of policies, budgets, committees, reports, or remediation projects do not by themselves demonstrate stronger safety constraints.

12. Anti-Double-Counting Rule

A single historical event may be relevant to more than one variable but must not automatically be treated as multiple independent pieces of evidence. Each overlapping entry must identify the shared source/event and state whether the second classification adds distinct information or is descriptive metadata only.

13. Reproducibility Test

A second scientifically literate evaluator should be able to apply this protocol to the same source set without access to the first evaluator's judgments. Agreement should be assessed for: variable applicability, direction, measure classification, and transferability outcome.

Disagreement is not automatically a protocol failure, but recurring disagreement caused by vague definitions, unstable denominators, or ambiguous rules is evidence that the measurement architecture needs revision after the frozen test is complete.

14. Dashboard-Level Pass/Fail Questions

·         Did the instrument distinguish technical vulnerability from demonstrated mitigation?

·         Did it register improving evidence before January 1, 2000 when warranted?

·         Did calendar proximity alone fail to drive the assessment upward?

·         Did it preserve uncertainty rather than forcing unsupported numerical precision?

·         Did it avoid treating planning, spending, or policy activity as demonstrated effectiveness?

·         Did it avoid obvious double counting?

·         Could another evaluator reproduce the classifications with reasonable agreement?

·         Did Variables 9 and 10 detect meaningful changes in demonstrated control and safety constraints?

·         Were AI-specific variables allowed to remain Not Applicable rather than being forced into the historical case?

·         Did the late-1999 dashboard state make sense when compared with the actual rollover outcome, without rewriting the pre-event record?

15. Overall Back-Test Outcome

The back-test does not receive a single arbitrary numerical score. Its outcome is a structured judgment based on transferability, measurability, reproducibility, directional sensitivity, and resistance to hindsight.

·         PASS: The dashboard tracks changing Y2K risk evidence coherently, especially demonstrated control improvement, without relying on hindsight or forced analogies.

·         PASS WITH AMENDMENTS: The architecture is useful but one or more definitions, denominators, or transfer rules require revision after defects are documented.

·         FAIL: The dashboard cannot distinguish meaningful changes in vulnerability/control, produces unstable or irreproducible classifications, or requires substantial retrofitting to fit Y2K.

16. Most Important Pre-Rollover Test

By late 1999, before the rollover occurred, would the frozen dashboard have recognized any demonstrated reduction in danger produced by remediation, testing, and strengthened safeguards?

If yes, the result supports the protocol's central rule that reality moves the assessment rather than time moving it automatically. If the dashboard only appears accurate after January 1, 2000, the historical back-test has provided little validation.

17. Amendment Rule

No defect found during the Y2K test may be silently repaired inside v1.0. Each defect is logged against the frozen version. After the complete back-test, proposed changes are issued as a later version with the original v1.0 preserved.

18. Version Record

Y2K Back-Test Protocol v1.0 - Frozen before historical evidence collection. Designed to test the previously frozen ten-variable dashboard by chronological replay, variable transfer analysis, and reproducibility.

Previous
Previous

AI Risk Dashboard v1 Frozen Reference

Next
Next

Y2K Back-Test Conclusion Statement