# Sealed Historical Challenge 01 — PROTOCOL v0.1

**Drafter:** grok · **Hub job:** `b73c292a-16e2-4f11-9370-1337d0239791`  
**Responds to:** James message `4b9818bb` (ASTRA — SRP SEALED HISTORICAL CHALLENGE 01)  
**Date:** 2026-09-10 ~1:55 PM PT  
**Status:** PROTOCOL_DESIGN_ONLY · **NOT frozen** · awaiting independent challenge (Claude) + Astra adjudication  
**Gates:** `MODEL_CHANGE=NO` throughout · no case selection · no outcome-window inspection · no threshold/model/score changes · no empirical SRP historical analysis

**Honesty label.** Design only. Do **not** execute until Astra freezes an accepted design. UNKNOWN ≠ 0 ≠ absence of risk.

---

## 0. Purpose / non-goals / authorization

**Core question.** Can the *currently* accepted/research-usable PARTIAL SRP measurements characterize historical U.S. societal states so as to distinguish (i) periods preceding substantial systemic stress from (ii) hard-negative periods that looked concerning but whose disturbances were subsequently absorbed — focusing on **system response / susceptibility**, not “predict crisis.”

**This is NOT:** final L9 validation · authorization to tune SRP · authorization to inspect historical outcomes yet.

### AUTHORIZED NOW
protocol design · source feasibility · integrity design · baseline design · scoring design · blinding design

### NOT AUTHORIZED (explicit)
- case selection / famous-crisis cherry-picking  
- future-outcome inspection / candidate outcome-window browsing  
- SRP historical analysis / scoring runs against sealed cases  
- threshold tuning · model changes · SRP scoring-rule changes  
- resurrecting Exp001 · Exp002 ratchet inference  
- generalizing Rigidity-01 beyond House SUPPORTED_DESCRIPTIVELY  
- RIGIDITY-02 mechanism claims  
- manufacturing values for unfinished constructs  
- production alerts / Phase-Load-Sync-PPS computation from partial SRP  
- claiming L9 validation from any exploratory result of this challenge

---

## 1. Blinding + agent separation

| Role | Who | Sees |
|------|-----|------|
| Protocol design / challenge | Grok (this doc) · Claude (independent challenge) · Astra (adjudicate) | Protocol text only — **no** case list |
| Case-selection / outcome coding | **Isolated selector** (separate process; not analysis agents) | Frozen outcome defs + PIT-safe source calendars; **not** SRP state assessments |
| Analysis (PARTIAL SRP + baselines) | Assigned analysis agents **after** protocol freeze + sealed case-bundle handoff | Cutoff-T PIT bundles **without** post-T outcomes, case narrative names that reveal futures, or outcome labels |
| PIT / leakage / bundle integrity | Codex | Bundle manifests, release≠fieldwork checks, leakage fixtures |
| Implementation after accept | Grok | Frozen protocol + integrity-cleared bundles only |
| Final adjudication | Astra | Locked pre-unblind outputs + unblinded coded outcomes |

**Hard rules.**
1. Case selection occurs **only after** protocol freeze.  
2. Analysis agents must not administer or preview the outcome-coding process.  
3. Selector must not see SRP assessments or baseline predictions.  
4. Case/outcome information separated from analyzing agents wherever technically possible (separate job artifacts, hashed sealed bundles, dual custody of unblind key).  
5. No agent may inspect candidate future outcome windows while drafting or revising this protocol.

---

## 2. Compact output contract (Step 1) — labels A–F measurability review

Labels are **proposals**. Prefer FEASIBLE / INFEASIBLE / UNKNOWN. Do not manufacture values.

| ID | Proposed labels | Measurability (v0.1) | May enter challenge output? | Reason (governance-tied) |
|----|-----------------|----------------------|-----------------------------|---------------------------|
| **A** | Institutional recovery capacity: NORMAL / IMPAIRED / UNKNOWN | **INFEASIBLE** as scored construct | Emit **UNKNOWN** only (optional L1 descriptive annex) | Recovery HL / institutional-recovery constructs remain UNKNOWN/INSUFFICIENT. L1-US-v0.1 is score-authorized for **its own construct only**, not recovery capacity. |
| **B** | Elite structural rigidity: LOW / MODERATE / HIGH / UNKNOWN | **INFEASIBLE** as ordinal score | Emit **UNKNOWN**; House descriptive flag allowed as *evidence note* only | Rigidity-01: House **SUPPORTED_DESCRIPTIVELY** only as currently accepted interpretation. Do **not** generalize Senate or claim causality. RIGIDITY-02 blocked. |
| **C** | Cross-system propagation susceptibility: LOW / MODERATE / HIGH / UNKNOWN | **INFEASIBLE** | **UNKNOWN** | Coupling / propagation constructs incomplete. Exp001 **CLOSED/INSUFFICIENT** — cannot resurrect. |
| **D** | Disturbance persistence susceptibility: LOW / MODERATE / HIGH / UNKNOWN | **INFEASIBLE** | **UNKNOWN** | Exp002 **INSUFFICIENT** — no ratchet inference. |
| **E** | Most vulnerable measured subsystem: \<construct\> / UNKNOWN | **PARTIAL / mostly UNKNOWN** | **UNKNOWN** unless a single allowlisted measured construct is the *only* non-UNKNOWN input (then name that construct with `partial_only=true`) | Cross-subsystem ranking requires unfinished constructs. S1/E2 exploratory parked — use only components whose status permits the exact quantity; else UNKNOWN. |
| **F** | Overall system-response: RESILIENT / MIXED / FRAGILE / UNKNOWN | **INFEASIBLE** as aggregate score | **UNKNOWN** (required default) | Resilience / Phase / Load / Sync / PPS / Alert = UNKNOWN/INSUFFICIENT. Must **not** be computed from L1 alone. |

**Frozen v0.1 machine output policy.** Primary challenge fields A–F default to `UNKNOWN` with explicit `reason_code` from allowlist (§6). Optional annex fields may carry L1-US-v0.1 band/status and Rigidity-01 House descriptive observation **without** promoting them into A–F scored labels.

**UNKNOWN policy.** UNKNOWN ≠ 0 ≠ “safe.” UNKNOWN does not authorize optimistic F=RESILIENT or A=NORMAL.

---

## 3. Outcome definitions (Step 2) — objective coding, no events/years

Horizon H = primary frozen horizon (§4). For each sealed cutoff \(T\), code the open interval \((T, T+H]\) using predeclared observable criteria. Coders must not use SRP outputs. Criteria use countable public series / institutional chronologies available under PIT rules; **no crisis names** in this protocol.

### ABSORBED
All of the following hold in \((T, T+H]\):
1. Any qualifying disturbance episode(s) that begin in-window (or that were already active at \(T\) and remain under watch) show **return-to-baseline** on the predeclared disturbance intensity index within H, **and**  
2. No secondary subsystem breach criteria (predeclared list) are met, **and**  
3. No SYSTEMIC_STRESS criteria (below) are met.

### PERSISTENT
At least one qualifying disturbance episode active in-window remains **above** its predeclared intensity threshold for a continuous duration ≥ \(d_P\) (freeze \(d_P\) before cases; proposal: 6 months) **without** meeting PROPAGATED or SYSTEMIC_STRESS.

### PROPAGATED
A qualifying disturbance originating in subsystem \(S_i\) is followed, within H, by breach criteria in ≥1 distinct subsystem \(S_j\) (j≠i) under the predeclared cross-subsystem edge list — without necessarily meeting SYSTEMIC_STRESS.

### SYSTEMIC_STRESS
≥2 of the following independent criteria hold within H (exact thresholds frozen before cases; listed here as **slots**, not tuned values):
- Multi-subsystem concurrent breach (≥k subsystems, k≥3 proposed)  
- Institutional continuity stress marker (predeclared countable rule: e.g., extraordinary continuity procedures triggered under public statute/record — **not** narrative labels)  
- Macro-real activity breach of predeclared magnitude/duration  
- Civic/security intensity index breach of predeclared magnitude/duration  
- Fiscal/financial stress marker from allowlisted public series breach

### AMBIGUOUS
Any of: insufficient PIT-safe outcome data; conflicting criteria across ABSORBED vs PERSISTENT; coder disagreement unresolved after adjudication protocol; horizon truncation (end of available PIT-safe record before \(T+H\)).

**Coding discipline.** Dual independent coders → adjudication sheet → lock. Disagreement ⇒ AMBIGUOUS unless adjudicator resolves under frozen tie-break rules. **Do not** inspect which historical periods satisfy these defs while revising them.

---

## 4. Horizon freeze (Step 3)

| Horizon | Status | Value |
|---------|--------|-------|
| **Primary** | **PROPOSAL pending Astra freeze** | **24 months** |
| Secondary | Optional; if used, freeze **before** cases | Proposal: 60 months (exploratory sensitivity only; not primary scoring) |

**Rules.** One primary horizon for all cases. No per-case horizon shopping. Secondary, if frozen, reported separately and cannot rescue primary failures.

---

## 5. Case-selection ALGORITHM (Step 4) — algorithm only; NO cases

**Target N:** 6–10 U.S. windows. Mix of **STRESS** and **HARD NEGATIVE** strata. Non-overlapping outcome windows. Outcome-independent / isolated selector.

### 5.1 Sampling frame
- Geography: United States (national-level coding unit).  
- Eligible cutoffs: month-end timestamps on a predeclared grid (proposal: calendar quarter-ends) spanning a predeclared long panel whose **start/end are fixed before selection** without reference to named episodes.  
- Each candidate cutoff \(T\) defines outcome window \(W(T)=(T, T+H]\).

### 5.2 Non-overlap
Reject any new \(T'\) if \(W(T')\) overlaps any already accepted \(W(T)\) by more than τ months (proposal τ=0: strict non-overlap of open intervals). Also enforce minimum gap \(g\) between cutoffs (proposal g=H).

### 5.3 Stratum definitions (outcome-driven; selector-only)
Uses **frozen outcome defs** only — not SRP:
- **STRESS pool:** cutoffs where coded outcome ∈ {SYSTEMIC_STRESS, PROPAGATED} (and optionally PERSISTENT if primary stress definition so assigns — freeze mapping before run).  
- **HARD NEGATIVE pool:** cutoffs where contemporaneous *concern screen* (below) is positive **and** coded outcome = ABSORBED.

### 5.4 Concern screen (outcome-independent features at T only)
Predeclare a small set of **PIT-available raw indicators** (not SRP scores) used solely to define “looked concerning” for hard-negative eligibility — e.g., allowlisted macro and institutional raw series crossing predeclared **untuned** screen thresholds set without reference to outcomes. Screen thresholds are frozen before pool construction; **forbidden:** tuning screens to known episodes.

### 5.5 Inclusion
1. \(T\) on eligible grid; full PIT source coverage for analysis allowlist at T (or explicit UNRESOLVED path).  
2. Outcome window fully codeable or marked AMBIGUOUS-eligible.  
3. Passes non-overlap.  
4. Assigned stratum membership under §5.3.

### 5.6 Exclusion
- Incomplete PIT coverage for required analysis inputs without UNRESOLVED handling.  
- Overlap violations.  
- Cutoffs whose public identity cannot be blinded (selector assigns opaque case_id; analysis sees `case_id` + T only).  
- Any cutoff introduced by human famous-episode nomination (**banned**).

### 5.7 Allocation algorithm (deterministic)
1. Build STRESS and HARD NEGATIVE pools via isolated selector.  
2. Target mix proposal: ⌈N/2⌉ STRESS, ⌊N/2⌋ HARD NEGATIVE (adjust if a pool is short; record shortfall → may yield INSUFFICIENT).  
3. Within each pool, sample without replacement using a frozen PRNG seed + rank key = hash(T ‖ stratum ‖ protocol_hash) ordered ascending; take first k.  
4. Publish sealed case-bundle manifest: `{case_id, T, stratum_label_sealed, source_manifest_hash}` to custody; stratum_label_sealed held from analysis agents until unblind.

### 5.8 Prohibitions
No James/Astra hand-picked famous cases. No browsing outcome windows by analysis agents. No post-hoc replacement of cases after seeing SRP outputs.

---

## 6. Point-in-time wall (Step 5)

For every cutoff T:
- SRP and baselines receive **only** information available by T.  
- Release date ≠ fieldwork date — use Method Library PIT guards (vintage/release ledgers; archive capture timestamps; `eligible_strict_PIT` when defined).  
- Ban post-T: polls, economic releases, event data, historical revisions, outcome labels, narratives, case names revealing future events.  
- Revised future data allowed **only** if explicitly labeled unavoidable sensitivity and segregated from primary.  
- When availability is uncertain → **UNRESOLVED** (not imputed; not treated as 0).

**Bundle requirement.** Each case bundle includes: cutoff T, PIT source manifest, fieldwork vs release audit table, hash of inputs, guard version ids.

---

## 7. Current-model allowlist (Step 6)

`MODEL_CHANGE=NO`. Only current governance-permitted quantities enter.

| Construct / artifact | Enter analysis? | Role | Notes |
|----------------------|-----------------|------|-------|
| **L1-US-v0.1** | YES | Score-authorized for **L1 construct only** | Frozen/live; bands per rule_id; missing→null/UNKNOWN |
| **M1** | NO as predictor | Construct-diagnostic only if needed for annex | `score_authorized=0`; UNKNOWN as predictor |
| **Rigidity-01 House descriptive** | YES as descriptive evidence note | SUPPORTED_DESCRIPTIVELY interpretation only | Do not emit B=LOW/MOD/HIGH from it; no Senate generalization; no causality |
| **RIGIDITY-02** | NO | Blocked | |
| **Exp001** | NO | CLOSED/INSUFFICIENT | Cannot resurrect |
| **Exp002** | NO | INSUFFICIENT | No ratchet |
| **S1 / E2** | CONDITIONAL | Only components whose **current** status permits the exact requested quantity | Else UNKNOWN; exploratory parked |
| **Phase / Load / Sync / PPS / Alert / Resilience / Coupling** | NO | UNKNOWN / INSUFFICIENT | Do not compute from L1 alone |
| Raw allowlisted PIT series for baselines B1/B2 | YES | Baseline inputs | Equal information vs SRP (§9) |

Unavailable ⇒ UNKNOWN.

---

## 8. Pre-unblind analysis output schema + lock (Step 7)

For every case, **before** unblinding, emit:

### 8.1 Human-readable packet
- STATE ASSESSMENT (PARTIAL SRP interpretation under allowlist; default UNKNOWN-heavy)  
- CONFIDENCE (separate axes: data_coverage, source_quality, construct_validity, freshness, overall — coverage ≠ confidence)  
- EVIDENCE USED  
- EVIDENCE AGAINST  
- UNKNOWN COMPONENTS (explicit list + reason_codes)  
- ALTERNATIVE INTERPRETATIONS  

### 8.2 Machine-readable prediction record (JSON)
```json
{
  "challenge_id": "sealed-historical-challenge-01",
  "protocol_version": "0.1",
  "case_id": "<opaque>",
  "cutoff_T": "<ISO-8601 date>",
  "horizon_months_primary": 24,
  "outputs": {
    "A": {"value": "UNKNOWN", "reason_code": "..."},
    "B": {"value": "UNKNOWN", "reason_code": "..."},
    "C": {"value": "UNKNOWN", "reason_code": "..."},
    "D": {"value": "UNKNOWN", "reason_code": "..."},
    "E": {"value": "UNKNOWN", "reason_code": "...", "construct": null},
    "F": {"value": "UNKNOWN", "reason_code": "..."}
  },
  "annex": {
    "L1_US_v0_1": {"band": null, "status": null},
    "rigidity01_house_descriptive": null
  },
  "baselines": {"B0": {}, "B1": {}, "B2": {}},
  "confidence": {},
  "evidence_used": [],
  "evidence_against": [],
  "unknown_components": [],
  "alternative_interpretations": [],
  "source_manifest_hash": "<sha256>",
  "agent_versions": {},
  "model_change": false
}
```

### 8.3 Lock requirements (before any future data revealed)
Lock and publish: full output artifact · UTC timestamp · sha256 · agent versions · source manifest · protocol_version hash. Custody holds unblind key separately. **No edits** post-lock except labeled errata that do not change predictions.

---

## 9. Baselines (Step 8) — equal information rule

| ID | Definition |
|----|------------|
| **B0** | Simple historical/base-rate assessment over the frozen sampling frame (stratum-blind marginal rates of outcome classes). |
| **B1** | Ordinary contemporaneous stress indicators from the **same PIT-available raw information** allowed to PARTIAL SRP (no SRP constructs). |
| **B2** | Ordinary temporal baseline: levels / trends / lags of the same accepted raw series (no SRP transforms). |
| **PARTIAL SRP** | Current system-state interpretation under §2+§7 allowlist. |

**Equal information rule.** No baseline may receive less information merely to make SRP look better. Conversely, baselines must not receive post-T or non-allowlisted leakage.

---

## 10. Scoring (Step 9) — N≈6–10; no fake significance

Predeclare before cases:

| Metric | Rule |
|--------|------|
| Accuracy | Case-level match of primary prediction class vs coded outcome mapping table (freeze mapping: e.g., F/UNKNOWN handling matrix before unblind). Given A–F mostly UNKNOWN under current governance, primary scored contrast may reduce to annex+baseline vs outcome strata — **state this limitation explicitly**; do not invent A–F values to create a score. |
| Calibration | Only if probabilities are emitted; otherwise N/A. |
| False-positive burden | Among HARD NEGATIVES, rate of non-UNKNOWN “fragile/stress-leaning” calls (if any emitted). |
| Hard-negative discrimination | Ability to avoid stress-leaning calls on ABSORBED hard negatives. |
| Stress-case discrimination | Ability to flag elevated susceptibility on STRESS strata **when** allowlist permits non-UNKNOWN signal; else report powerlessness honestly. |
| UNKNOWN handling | Track UNKNOWN rate; high UNKNOWN is a valid result (INSUFFICIENT informativeness), not a silent pass. |
| Confidence scoring | Compare stated confidence vs correctness; punish overconfidence on UNKNOWN-heavy packets. |
| Case-level score | Primary unit of evidence (per-case card). |
| Aggregate score | Descriptive summary only (mean/median of case scores; bootstrap CIs optional). **Do not** claim statistical significance from N≈6–10. |

**Mapping freeze.** Before unblind, freeze how PARTIAL SRP outputs (including all-UNKNOWN) translate into comparable predictions vs B0/B1/B2 for the outcome classes.

---

## 11. Failure / result enums (Step 10)

| Enum | Meaning | Authorizes | Does **not** authorize |
|------|---------|------------|------------------------|
| **SUPPORTED_EXPLORATORILY** | Predeclared metrics favor PARTIAL SRP vs baselines on sealed set | Further research design / replication proposals | SRP validation · scoring authorization · criticality/causality proofs · L9 replacement · production alerts |
| **MIXED** | Split case-level pattern / metric conflict | Targeted follow-up protocols | Same as above |
| **NOT_SUPPORTED** | PARTIAL SRP fails to beat baselines / wrong-direction under allowlist | Reporting negative exploratory evidence on **current operationalization** | Automatic falsification of every SRP theory |
| **INSUFFICIENT** | Too many UNKNOWN / too little allowlisted signal / pool shortfall / PIT failure | Measurement roadmap pressure | Claims that “no risk” or that SRP “works” |

Negative result evaluates the **CURRENT PARTIAL operationalization**, not the entire theoretical program.

---

## 12. Independence matrix

| Agent | Duty |
|-------|------|
| **Claude** | Independent protocol/method challenge (attack blinding holes, outcome circularity, allowlist overreach, scoring inflated claims) |
| **Codex** | PIT integrity · leakage fixtures · case-bundle integrity |
| **Grok** | Implementation **only after** protocol acceptance |
| **Astra** | Final protocol adjudication and synthesis |

---

## 13. Freeze checklist (before any case work)

- [ ] Claude challenge filed and addressed or explicitly waived by Astra  
- [ ] Astra ACCEPT (or CONDITIONAL with listed patches applied)  
- [ ] Primary horizon frozen (24 months proposal)  
- [ ] Secondary horizon accept/reject frozen  
- [ ] Outcome defs + thresholds slots filled with numeric freezes  
- [ ] Concern-screen series + untuned thresholds frozen  
- [ ] Non-overlap τ, gap g, N, stratum mix frozen  
- [ ] PRNG seed custody established  
- [ ] Allowlist table unchanged (`MODEL_CHANGE=NO`)  
- [ ] Scoring mapping + UNKNOWN policy frozen  
- [ ] Output schema + lock procedure rehearsed on synthetic empty bundle  

**STOP.** Do not select cases or run SRP historical analysis until the above checklist is complete under Astra freeze.

---

## 14. Document control

| Field | Value |
|-------|-------|
| protocol_id | `sealed-historical-challenge-01` |
| version | `0.1` |
| hub_job | `b73c292a-16e2-4f11-9370-1337d0239791` |
| james_message | `4b9818bb` |
| local_md | `/workspace/srp-observatory/docs/sealed-historical-challenge-01/PROTOCOL_v0_1.md` |
| local_json | `/workspace/srp-observatory/docs/sealed-historical-challenge-01/protocol_v0_1.json` |
| hub_slug_md | `/docs/sealed-historical-challenge-01-protocol` |
| hub_slug_json | `/docs/sealed-historical-challenge-01-protocol.json` |
