# ELECTION-01 integrity and acceptance contract

James explicitly requested that the team not cheat. These are required acceptance conditions, including protection against accidental leakage and selective reporting. This document defines review gates; it is not evidence that the Hub or backtest already enforces them in code.

1. Before fitting, publish an independently reviewed, hashed protocol: target, cutoff/timezone, eligible cycles, feature registry, sources/vintages, baseline, model, tuning budget, folds, seeds, metrics, practical success threshold and stopping rules. Coverage-led changes before outcome inspection must be versioned. Never silently overwrite a plan.
2. Keep a ledger of every attempted variant, failed run, exclusion and amendment. Exploratory findings stay exploratory. After inspecting test performance, a change consumes that test set for model selection; rerunning on it does not restore an independent test. Preserve the original result and seek genuinely unexamined evaluation or a frozen prospective forecast.
3. Require availability evidence for every predictor and every revised vintage at each cutoff. Fieldwork dates alone are insufficient. Exclude unknown availability from the strict backtest. Reject future pollster ratings, outcome-derived features, post-election releases and later revisions.
4. Separate fitting and evaluation interfaces. Held-out outcomes must be unavailable to fitting, feature selection, preprocessing and tuning. All transformations are learned from past training cycles. Freeze and hash predictions before the evaluator joins test labels. Hashes document artifacts; they do not prove absence of prior knowledge or tampering without a trustworthy record.
5. Required executable negative controls: future-release records are rejected; unavailable revisions are rejected; changing held-out labels cannot alter predictions or feature selection; no election crosses training/test boundaries; overlapping samples are not counted as independent polls; comparisons use identical eligible test cycles. Attach actual outputs, not claims that tests should pass.
6. Report every test election's prediction/error, all exclusions and missingness, and all attempted models. Report uncertainty at the election level. Do not manufacture sample size from poll rows or replicated national values. Baselines receive fair, predeclared tuning budgets; no selectively weak comparator.
7. A separate reviewer must inspect leakage and reproduce predictions/metrics from the locked artifacts. Require raw-source hashes, provenance, code/version, dependency environment, configuration and run logs. Distinguish code rerun from independent source replication. The proposer cannot give sole final approval or self-certify REPLICATED.
8. If provenance, timing or reproducibility fails, mark the result BLOCKED or exploratory, stating the failed gate. No invented values, concealed failures, relabeled retrospective results or claims of novel predictive power. Preserve null/negative findings. Historical outcomes are publicly known: even a careful retrospective test is not a genuinely blind prospective test.

Grok owns source ledger/run artifacts; Claude independently challenges the protocol and reproduction; Codex verifies the acceptance evidence and reports limitations. Reviewer reassignment must preserve independence. Existing SRP scoring/model authorizations remain untouched. No additional permission from James is required for these routine checks.
