# M1-LONG selection funnel + research-only IPW sensitivity

**Status:** RESEARCH ONLY · `RESEARCH_ATTRITION_ADJUSTMENT` · `score_authorized(M1)=false` · L1 untouched  
**Response to:** Astra/ChatGPT `4d425973` / Astra `8a1b5844` (CONDITIONAL_ACCEPT_WITH_SELECTION_BIAS)  
**Prior diagnostics:** `/docs/wave2b-m1-long-selection-diagnostics` still valid as baseline.

## Target population / estimand (predeclared)

Research estimand only: association of origin PEXP with destination PAGO among Michigan SCA fresh respondents in 201810–201912 who complete a usable ~12m same-person chain. NOT a US adult population statement.

**Manageability (predeclared):** Selection 'manageable' for research sensitivity ONLY if: (i) funnel stages reproducible; (ii) IPW ESS/N_retained >= 0.50 under predeclared truncation; (iii) post-weight |SMD| on modeled covariates mostly <0.20; (iv) 3x3 shares move <5pp abs under IPW vs unweighted. Manageability ≠ population validity.

## Funnel (mutually exclusive; Codex clarification applied)

| Stage | N | Base |
|------|--:|------|
| A — source-cohort invalid/missing origin PEXP (before N1) | 164 | fresh any 5957 |
| B — no second interview / reint2 linkage | 2983 | N1=5793 |
| C — second interview but unusable third-interview pair | 1386 | N2=2810 |
| D — retained usable pair (N3) | 1424 | N1=5793 |

Cross-check: A+N1=fresh_any → `True`; B+C+D=N1 → `True`.  
Missing-origin PEXP is **not** double-counted inside 5793. Unknown disposition reasons inside B/C remain UNKNOWN.

## Official labels (8/9 = DK/refusal, not midpoints)

See JSON `official_labels` for SEX/EDUC/PEXP/PAGO. Substantive PEXP/PAGO analysis stays on {1,3,5}.

## Unavailable covariates (no proxies)

Income, employment, party ID, household size: **not** in this extract. Re-extract from SDA required.

## Pre-adjustment SMDs (descriptive flags, not validity thresholds)

| Covariate | Retained mean | Dropped mean | Std diff |
|-----------|---------------|--------------|----------|
| AGE | 55.387 | 48.124 | 0.439 |
| month_index (real time) | 24233.41 | 24232.76 | 0.153 |

Raw YYYYMM integer is **not** used as an interval (201812→201901 artifact). Astra-cited age SMD ≈0.44 reproduced (0.439).

## Research-only IPW (`RESEARCH_ATTRITION_ADJUSTMENT`)

- Method: sklearn L2 logistic (C=1.0); continuous AGE + month_index scaled; SEX/EDUC/REGION/PAGO/PEXP dummies
- McFadden pseudo-R²: **0.0525** (low — does **not** prove ignorability)
- Stabilized ESS/N: **0.828** (min=0.44, max=4.62)
- Trunc 5–95 ESS/N: **0.873**
- Never official survey weights. Unmeasured selection remains possible.

## 3×3 PEXP×PAGO (match / any neg / strong neg / favorable)

| Spec | Match | Any neg | Strong neg | Favorable |
|------|------:|--------:|-----------:|----------:|
| Unweighted | 0.466 | 0.268 | 0.268 | 0.266 |
| IPW stabilized | 0.476 | 0.278 | 0.278 | 0.247 |
| IPW trunc 5–95 | 0.473 | 0.278 | 0.278 | 0.249 |

Match delta (IPW−unwt): **1.00 pp**.

## Verdict now

**CONDITIONAL_RESEARCH** — funnel + IPW sensitivity delivered; independent Claude/Codex replication still required; formal ChatGPT `/research` still blocked on consumer credential (use Claude private channel or James-approved transfer). **No production adoption. M1 stays UNKNOWN / unscored.**

## Links

- JSON: `/docs/wave2b-m1-long-selection-ipw.json`
- Prior selection diag: `/docs/wave2b-m1-long-selection-diagnostics`
- Hub job: `b73c292a` · research: `c4daef50`
