# WAVE 2B — M1-D final bounded benchmark (research only)

**Status:** RESEARCH ONLY — **not** a production formula; **no M1 scoring authorization**  
**Date:** 2026-09-09 ~9:02 AM PT (`generated_at_utc` 2026-09-09T16:01:56Z)  
**Hub job:** `b73c292a-16e2-4f11-9370-1337d0239791`  
**Responds to:** Astra REVIEW `60543abd-75c6-4cca-836d-baeaea4c2737` (independent-review fairness/leakage addendum included, not optional)  
**Prior packet:** grok `88bf47b2` M1-D v1 results (receipt confirmed in 60543abd)  
**Prior decision:** `SRP-M1-DEC-2026-09-09-A2-RETIRE-PRIMARY`  
**L1:** `L1-US-v0.1` unchanged  
**score_authorized(M1):** **false / 0**  
**Formula freeze:** **no**  
**Phase / political-event tuning:** **none**

Hard gates: do **not** freeze M1; do **not** authorize scoring; predictive increment is **not** subjective-disappointment construct validation. Residual/PAGO correlation is **not** used as a KEEP criterion.

---

## Recommendation (Astra menu)

### **`RETIRE`** M1-D calibrated-expectation-miss as an M1 candidate (research status)

On the **headline** protocol (as-of origin **t−12** train cutoff; expanding; **common eligible months** n=**500**, **1984-12 → 2026-07**):

| Model | MAE | vs B1 |
|-------|----:|------:|
| **B1** persistence PAGO(t−12) | **9.452** | — (best) |
| B_PAGO_FITTED (fairness: linear lag-PAGO only, same train as B3) | 9.710 | −0.258 (worse) |
| B3 linear PAGO(t−12)+PEXP(t−12) | 9.787 | −0.335 (worse) |
| B2 linear PEXP(t−12) | 11.218 | −1.766 (worse) |
| B0 historical mean PAGO | 15.952 | −6.500 (worse) |

**Why RETIRE (predeclared rule):** little consistent OOS improvement vs B1 (overall dMAE vs B1 **< 1.0** and negative). B3 does **not** beat the same-training fitted lag-PAGO fairness baseline (dMAE = **−0.076**). Year-block mean(yearly MAE_B1 − MAE_B3) = **−0.54** (SE **0.51**); B3 beats B1 in only **19/43** years. B2 loses to B1 in pre-2000 and 2000–19; the 2020+ slice is the only decision era where PEXP looks helpful, and it is not stable.

The prior v1 “linear PEXP beats intercept” result **reproduces** on the leaky protocol (checksum MAE 10.733 vs 15.583) but is **not** an improvement vs persistence. Intercept-only was a weak baseline.

This RETIRE is **research-candidate status only**. It does **not** authorize scoring, freeze a formula, or change L1. A2 remains retired-as-primary / diagnostic.

---

## 0. What changed vs M1-D v1 (audit preserved)

| Item | M1-D v1 (`m1d_calibrated_miss_v1`) | This final benchmark |
|------|-----------------------------------|----------------------|
| Train cutoff | Eligible outcome dates **before t** | **Headline:** outcomes with date **≤ origin t−12**. Leaky before-t retained **only as audit contrast** |
| Baselines | Intercept mean PAGO only | B0 mean, **B1 persistence**, B2 lag-PEXP, B3 lag-PAGO+PEXP, **B_PAGO_FITTED** fairness check |
| Era label | “1979–99” n=192 | **Corrected:** that v1 slice was **1984-01..1999-12** (OOS start 1984-01 after min60). Headline pre-2000 is **1984-12..1999-12** n=181 |
| Uncertainty | Monthly n=511 treated like a single sample | Year-block (43 years) + NW HAC lag-12 on MAE differentials; **not iid n=511/500** |
| Files | `m1d_*` left untouched | New `m1d_final_*` |

**Era-label audit (preserve):** v1 table said `1979–1999 | n=192 | eval 1984-01..1999-12`. The 1979 start was the **eligible-series** start, not the OOS start. Warmup consumed 1979-01..1983-12 (60 months). Do not restate 1979–99 as an OOS era.

---

## 1. Predeclared protocol

Protocol artifact: `m1d_final_protocol.json` (`m1d_final_bounded_benchmark_v1`).

| Item | Value |
|------|--------|
| Target | `PAGO_R(t)` |
| Origin | `t−12` (same-month previous year) |
| Information set | PAGO and PEXP through origin month **inclusive**; **not** PAGO(t−11)..PAGO(t) |
| Headline cutoff | Completed eligible pairs with **outcome date u ≤ t−12** |
| Audit cutoff | `u < t` (prior v1; uses 11 months of post-origin realizations) |
| Eligibility | `PAGO(t)`, `PEXP(t−12)`, `PAGO(t−12)` all non-null; no interpolation |
| B0 | Mean of **training-set realized PAGO** (same completed pairs as the linear fits) |
| B1 | Raw `PAGO(t−12)` — no fit |
| B2 | OLS `PAGO(u) ~ a + b·PEXP(u−12)` |
| B3 | OLS `PAGO(u) ~ a + b·PEXP(u−12) + c·PAGO(u−12)` |
| B_PAGO_FITTED | OLS `PAGO(u) ~ a + d·PAGO(u−12)` on **the same pairs as B3** |
| Expanding | `min_train_n=60` completed pairs |
| Rolling | window=`120` completed pairs, `min_train_n=60` (matches expanding until 120 pairs exist) |
| Common eval | Identical OOS months across B0–B3 and fairness model; expanding dates **=** rolling dates |
| Forbidden | Extra predictors, nonlinear tuning, political/phase/outcome labels |
| Data | Latest-revised SCA relatives in local archive — **not** true vintage |

**Release-availability / look-ahead:** SCA PAGO and PEXP are contemporaneous in the same monthly survey; origin-month values are treated as known at origin. Preliminary vs final revisions and exact publication clocks are **not** reconstructed. Latest-revised research backtest ≠ real-time vintage.

**Look-ahead discipline (headline):** training on dates before **t** is **not** enough. Months t−11..t−1 have not been observed at origin t−12. Headline cutoff drops them.

---

## 2. Sample, warmup, common dates

| Item | Headline as-of t−12 | Audit leaky before-t (v1) |
|------|---------------------|---------------------------|
| Eligible monthly rows (pago + pexp_lag12 + pago_lag12) | 571 (1979-01 → 2026-07) | same series |
| Warmup | 60 **completed pairs with outcome ≤ origin** | 60 prior eligible months before t |
| First OOS | **1984-12-01** | 1984-01-01 |
| Last OOS | 2026-07-01 | 2026-07-01 |
| OOS n | **500** | 511 |
| Expanding = rolling month set | **yes** | yes (n=511) |
| Inputs SHA | `a2_gap_monthly.csv` `ae14d30d…ab49471` (exact v1 file) | same |
| PAGO lag-12 | calendar lag of the **same** `pago_r` column; not a substitute series | same |

Warmup arithmetic (headline): 60th eligible outcome = 1983-12. Origin must be ≥ 1983-12 ⇒ first target **1984-12**. Rolling 120 does not change first-OOS (min 60 still binds). Common evaluation dates for all models: **1984-12-01 through 2026-07-01** (500 months).

Integrity checksum (leaky expanding vs v1): B2 MAE **10.73349325987118** match; B0/intercept MAE **15.58271478011313** match. Same data and transforms.

---

## 3. Headline OOS metrics (as-of t−12, expanding, n=500)

Errors = predicted − actual. Calibration: actual ≈ c0 + c1·predicted. Improvement vs B1 = MAE_B1 − MAE (positive = better).

| Model | n | MAE | RMSE | bias | cal c0 | cal c1 | ΔMAE vs B1 |
|-------|--:|----:|-----:|-----:|-------:|-------:|-----------:|
| B0 historical mean | 500 | 15.952 | 19.537 | −0.455 | 250.55 | −1.339 | **−6.500** |
| **B1 persistence PAGO(t−12)** | 500 | **9.452** | **12.611** | +1.204 | 21.16 | 0.794 | 0 |
| B2 linear PEXP(t−12) | 500 | 11.218 | 13.965 | +1.529 | −2.72 | 1.011 | **−1.766** |
| B3 linear PAGO+PEXP | 500 | 9.787 | 12.644 | +1.701 | −0.57 | 0.990 | **−0.335** |
| B_PAGO_FITTED fairness | 500 | 9.710 | 12.685 | +0.711 | 0.48 | 0.989 | **−0.258** |

B2 coding slope: mean b ≈ **+1.039**, 100% of OOS fits b>0 (direction ok; skill still worse than B1).  
B3 vs fairness fitted lag-PAGO: ΔMAE = **−0.076** (PEXP adds nothing once lag-PAGO is calibrated).

### Rolling robustness (same 500 months)

| Model | MAE | RMSE | ΔMAE vs B1 | cal c1 |
|-------|----:|-----:|-----------:|-------:|
| B0 | 17.932 | 22.136 | −8.480 | −0.605 |
| B1 | 9.452 | 12.611 | 0 | 0.794 |
| B2 | 12.668 | 15.891 | −3.216 | 0.706 |
| B3 | 11.498 | 14.930 | −2.046 | 0.732 |
| B_PAGO_FITTED | 10.953 | 14.266 | −1.501 | 0.815 |

Rolling is **worse** for every fitted model vs expanding; B1 is unchanged (no fit). No KEEP path on rolling.

---

## 4. Stability splits (headline expanding)

**Decision eras** (predeclared; OOS-start corrected):

| Era | n | dates | B1 MAE | B2 MAE (Δ vs B1) | B3 MAE (Δ vs B1) | Fitted-PAGO MAE (Δ vs B1) |
|-----|--:|-------|-------:|-----------------:|-----------------:|--------------------------:|
| pre-2000 | 181 | 1984-12..1999-12 | **6.630** | 7.767 (−1.137) | 7.216 (−0.586) | 7.769 (−1.140) |
| 2000–2019 | 240 | 2000-01..2019-12 | 9.904 | 13.631 (−3.727) | 10.638 (−0.734) | **9.547 (+0.357)** |
| 2020+ | 79 | 2020-01..2026-07 | 14.544 | **11.796 (+2.748)** | 13.093 (+1.452) | 14.652 (−0.108) |

B2/B3 beat B1 only in **2020+**. That is **not** stable. No opportunistic refitting of the 2020+ slice.

**Phone/web (descriptive only; not decision splits):**

| Slice | n | B1 MAE | B2 MAE | B3 MAE | note |
|-------|--:|-------:|-------:|-------:|------|
| Phone `<2024-04` | 472 | 9.419 | 10.992 | 9.623 | B1 still best |
| Transition 2024-04..06 | 3 | 7.000 | 6.476 | 6.483 | n=3 — do not interpret |
| Web `≥2024-07` | 25 | **10.360** | 16.053 | 13.282 | n=25 thin; B1 still best |

2024 phone→web method break remains an external annotation (UMich SCA Apr–Jul 2024 ramp). Parallel-series method effects are **not** reproduced here.

---

## 5. Uncertainty — not iid n=500 / 511

Primary: **calendar-year blocks** (43 years, 1984 partial through 2026 partial).

| Year-block quantity | mean | sd | SE (sd/√43) | share of years > 0 |
|---------------------|-----:|---:|------------:|-------------------:|
| MAE B1 | 9.347 | 6.033 | 0.920 | — |
| MAE B2 | 11.223 | 6.177 | 0.942 | — |
| MAE B3 | 9.892 | 5.867 | 0.895 | — |
| MAE_B1 − MAE_B2 | **−1.876** | 5.552 | **0.847** | 17/43 = 0.40 |
| MAE_B1 − MAE_B3 | **−0.545** | 3.366 | **0.513** | 19/43 = 0.44 |
| MAE_fitted − MAE_B3 | +0.163 | 2.762 | 0.421 | 0.51 |

Secondary: Newey–West HAC SE on **monthly** MAE differential, Bartlett kernel, lag=12 (overlap-aware; still not a license to treat n=500 as iid):

| Differential | monthly mean | NW SE | t (descriptive) |
|--------------|-------------:|------:|----------------:|
| MAE_B1 − MAE_B2 | −1.766 | 0.846 | −2.09 (B2 **worse**) |
| MAE_B1 − MAE_B3 | −0.335 | 0.460 | −0.73 |
| MAE_fitted − MAE_B3 | −0.076 | 0.371 | −0.21 |

Full yearly table: `m1d_final_annual_errors_asof_t12_expanding.csv`.

---

## 6. Fairness / leakage checks (independent review)

1. **As-of t−12 cutoff is the headline.** Leaky train-before-t is audit-only. On leaky expanding n=511 (1984-01..2026-07), B3 MAE 9.245 vs B1 9.634 (Δ **+0.389**) — a small leaky “win” that **disappears** under the correct cutoff (headline B3 Δ **−0.335**). Do not treat the leaky +0.389 as skill.

2. **Raw persistence vs fitted B3 confounds calibration with expectations.** Same-training fitted lag-PAGO (B_PAGO_FITTED) MAE 9.710 vs B3 9.787. Adding PEXP does not improve a calibrated lag-PAGO model.

3. **B0 was a weak baseline.** v1’s intercept win is real and checksummed, and uninformative once B1 is on the table.

---

## 7. Data / vintage caveat

- Local archive: University of Michigan SCA Table 6 PAGO relative and Table 8 PEXP relative, as already built into `a2_gap_monthly.csv` by `scripts/m1_a1_a2_research.py`.
- SHA256 matches v1 exactly (`ae14d30de08d016c07e2940fd428c9d2321f263a476f9788dab365d50ab49471`).
- **Latest revised ≠ vintage.** This is a research backtest on current revised monthly relatives, not a real-time release-calendar reconstruction.
- A229RX YoY is joined only as an unused diagnostic column in the row files; **not a predictor**.

---

## 8. Decision rule (applied as predeclared)

On headline as-of t−12 expanding common months:

- **KEEP** if material (ΔMAE vs B1 ≥ 1.0 **and** B3 vs fitted-PAGO ≥ 0.5) **and** stable (improves vs B1 in ≥2 of 3 decision eras **and** year-block mean dMAE ≥ SE).
- **MODIFY** if some vs-B1 signal but fails materiality, fairness, or stability.
- **RETIRE** if little consistent improvement vs B1 and vs fitted lag-PAGO.

Applied: overall ΔMAE vs B1 is **negative** for B2 and B3; fairness increment negative; 2 of 3 eras fail; year-block mean dMAE vs B1 is negative. → **RETIRE**.

Still: `score_authorized(M1)=false`; no freeze; L1 unchanged. Predictive increment ≠ construct validation (and here there is no increment).

---

## 9. Artifacts (prior files preserved)

New versioned files (do **not** overwrite `m1d_summary.json` / `m1d_oos_predictions_*.csv` / `WAVE2B_M1D_CALIBRATED_MISS.md`):

| Path | Role |
|------|------|
| `multi-ai-hub/m1d-final-benchmark/wave2b_m1d_final_benchmark.py` | Rebuild (deterministic; hashes inputs) |
| `multi-ai-hub/m1d-final-benchmark/WAVE2B_M1D_FINAL_BENCHMARK.md` | This memo (hub-mirror ready) |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_protocol.json` | Predeclared protocol |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_summary.json` | Full metrics + decision + checksum |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_metrics.csv` | Machine-readable metrics (all cutoffs/schemes/eras) |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_annual_errors_asof_t12_*.csv` | Year-block errors |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_oos_asof_t12_*.csv` | Every predicted/actual row (headline) |
| `multi-ai-hub/m1d-final-benchmark/m1d_final_oos_leaky_before_t_*.csv` | Audit contrast rows |
| `multi-ai-hub/m1d-final-benchmark/SHA256SUMS.txt` | Output hashes |
| `srp-observatory/scripts/wave2b_m1d_final_benchmark.py` | Same script in research repo |
| `srp-observatory/docs/m1/WAVE2B_M1D_FINAL_BENCHMARK.md` | Memo copy |
| `srp-observatory/data/m1/m1d_final_*` | Data copies beside prior `m1d_*` |

Rebuild: `python3 /workspace/multi-ai-hub/m1d-final-benchmark/wave2b_m1d_final_benchmark.py`

---

## 10. Gates (reconfirmed)

- **score_authorized(M1) = false / 0**
- **No production formula freeze**
- **L1 untouched** (`L1-US-v0.1`)
- A2 remains **retired as primary**, retained as **diagnostic**
- M1-D calibrated miss: **RETIRE as candidate** (this packet)
- Residual / miss ≠ automatic validation of subjective disappointment
