Result RETAINED BASELINE RECORDED NULL
EP/WP tournament v1 — challengers vs. the nflverse baseline #
Gridiron Signal · NFL · Disposition 2026-08-05
Question Can any independent expected-points or win-probability arm beat replaying the nflverse EP/WP columns under frozen gates?
What happened
“EP/WP tournament v1: RETAIN_BASELINE (NULL result, recorded) … A NULL result is a publishable outcome, not a failure to be
massaged.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters Per the packet's decision field: "no independent arm beat source replay on the frozen gates." Production displays source EP/WP with nflverse attribution; no independent EP/WP model is claimed. The 2026-08-08 independent audit retained the disposition — challenger_status "REWORK", current_verdict "RETAIN_BASELINE" — pending a new predeclared future-data tournament.
Technical record
Result RETAINED BASELINE
QB CORE pilot v1 — context-adjusted quarterback impact #
Gridiron Signal · NFL · Disposition 2026-08-08
Question Does a context-ridge QB impact estimate earn research status against frozen gates?
What happened
“RETAIN_BASELINE … All three audited candidate-minus-raw effects are small and their player-cluster 95 percent intervals span zero; alpha 1000 also fails the split-half comparison with raw EPA/dropback.”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters The builder’s D-011 promotion is superseded: D-019 states "D-011 does not promote `passing_impact_rate`." The metric is scaffold, candidate use is research-workbench only, and no leaderboard or public values are authorized.
Technical record
SRC gridironsignal.rodericrinehart.com/data/experiments.json (qb_core_pilot_v1.independent_audit, audit_id qb_core_pilot_v1_codex_hostile_audit_2026-08-08) + Gridiron Signal DECISIONS.md D-019 (repo doc, not public) methodology — QB CORE pilot v1 — context-adjusted quarterback impact ↗
Result WITHHELD
Trench visibility v1 — can offensive linemen be identified at all? #
Gridiron Signal · NFL · Disposition 2026-08-08
Question Is individual offensive-line attribution identifiable enough to publish research values?
What happened
“WITHHOLD_INDIVIDUAL … quantitative results mix participation seasons withheld by policy”
Decision Keep the research. Do not publish.
Why it matters D-019: "D-013 does not permit individual offensive-line values. Current disposition is `WITHHOLD_INDIVIDUAL`; only unit and identifiability research is authorized." The identifiability measurements stand as recorded history; the publication permission does not.
Technical record
Result WITHHELD
Trench wave 2 — interior defensive line and off-ball linebackers #
Gridiron Signal · NFL · Disposition 2026-08-08
Question Do defensive-front research values survive a cross-family agreement gate?
What happened
“UNIT_IDENTIFIABILITY_ONLY”
Decision Keep the research. Do not publish.
Why it matters D-019: "D-015 does not permit individual defensive-front values. Current disposition is `UNIT_IDENTIFIABILITY_ONLY`; missing assignment evidence remains unmeasured." The family-spread finding (mean cross-family agreement 0.544) remains recorded history.
Technical record
SRC gridironsignal.rodericrinehart.com/data/experiments.json (trench_visibility_v2_front.audit_posture, publication_status "withheld") + Gridiron Signal DECISIONS.md D-019 (repo doc, not public) methodology — Trench wave 2 — interior defensive line and off-ball linebackers ↗
Result QUARANTINED
Season conservation under a hostile audit — 22 of 27 PASS #
Gridiron Signal · NFL · Correction 2026-08-08
Question Which historical seasons survive validation against source play-by-play defects?
What happened
“The former 24/27 conservation claim did not check reverse schedule coverage or fail
on errors. The repaired v2 result is 22/27 PASS.”
Decision Isolate from publication. Keep the disposition as history.
Why it matters Five seasons are publication-ineligible under conservation-v3 — 2001 and 2002 each contain 8 fatal games (transposed scores in source play-by-play), 2011 has one phantom-TD-class fatal game, and 1999–2000 fail the repaired reverse-coverage checks. Quarantined seasons stay quarantined rather than being repaired by guesswork.
Technical record
Result RETAINED BASELINE
PULSE EWMA challenger — exponential decay vs. transparent windows #
Hardball Signal · MLB · Disposition 2026-08-06
Question Does an EWMA form estimate beat transparent trailing windows?
What happened
“Transparent windows retained (EWMA lost)”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters PULSE is descriptive-only, never a forecast. The v1 forward-test numbers once quoted here were superseded when the 2026-08-08 post-moonshot audit found the test had leaked target information (PM-01); the corrected, sealed v2 evidence retains the anti-forecast finding — see the forward-test correction entry.
Technical record
SRC Hardball Signal ROADMAP.md at commit cc8854e (historical MS-13 row; repo doc, not public) methodology — PULSE EWMA challenger — exponential decay vs. transparent windows ↗
Result FAILED
Park-adjusted batting tournament: challenger failed, result published #
Hardball Signal · MLB · Disposition 2026-08-06
Question Does park adjustment improve the batting impact-rate model enough to promote?
What happened
“park arm FAILED (-2.3e-05 vs 0.001 threshold) → incumbent retained, failed arm published”
Decision Keep the failure on the record. Do not publish this approach as a success.
Why it matters A reformulated additive arm later landed at -0.000305 — "still a null, but a real one," with power proven on synthetic data.
Technical record
SRC Hardball Signal ROADMAP.md A2 row + docs/operations/DEPLOYMENT.md 3d6eb6d0 row (repo docs, not public) validation — Park-adjusted batting tournament: challenger failed, result published ↗
Result RETAINED BASELINE
Starter and reliever pitching tournaments — claims dissolved #
Hardball Signal · MLB · Disposition 2026-08-06
Question Do the pitching impact-rate challengers hold up under paired bootstrap?
What happened
“starter 95% [-0.003991, +0.000468] and reliever [-0.001910, +0.000483] both SPAN ZERO and are now labelled not_distinguishable_from_zero”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters "Green" was demoted to meaning the tournament executed soundly — not that a claim holds.
Technical record
SRC Hardball Signal docs/operations/DEPLOYMENT.md, deployment f54f3f56 row (repo doc, not public) validation — Starter and reliever pitching tournaments — claims dissolved ↗
Result BLOCKED BY RIGHTS
Statcast-derived metrics — complete, tested, and blocked #
Hardball Signal · MLB · Disposition 2026-08-07
Question Can pitch-level Statcast metrics ship?
What happened
“The seam is complete and tested; **no admitted lawful source exists**”
Decision Stop until a lawful source is admitted.
Why it matters The product's /pitches page serves an honest unavailable state with a named unblocking condition rather than scraping.
Technical record
SRC Hardball Signal docs/releases/FINAL_REPORT_2026-08-07.md line 72 (repo doc, not public) data & sources — Statcast-derived metrics — complete, tested, and blocked ↗
Result RESEARCH RECORDED NULL
Lagged CORE vs. public metrics on a matched future-outcome task #
Court Signal · NBA · Disposition 2026-07-29
Question Does lagged CORE beat BPM, PER, WS/48, and VORP-rate at a frozen prediction task?
What happened
“Benchmarks pilot complete — verdict NULL, 8/16 criteria, no promotion.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters "No completed diagnostic or experiment nominates a production formula change." CORE remains provisional by its own registry.
Technical record
SRC Court Signal ROADMAP.md header + docs/ANALYTICS_EVIDENCE_CHECKPOINT_2026-07-29.md, receipt ae454ca9… (repo docs, not public) evidence — Lagged CORE vs. public metrics on a matched future-outcome task ↗
Result BLOCKED BY RIGHTS
Shot charts and spatial analytics — parked on permission, not failure #
Court Signal · NBA · Disposition 2026-08-01
Question Can shot-coordinate spatial surfaces ship?
What happened
“parked external dependency, not a
failed product”
Decision Stop until a lawful source is admitted.
Why it matters v3.1.2 tightened the epistemic wording: "No source is admitted. The request is reported sent by Roderic and awaiting a response; that is not permission or proof of delivery." Remaining at Tier 0 permanently is declared an acceptable final state.
Technical record
SRC Court Signal ROADMAP.md CONTROLLING POSTURE + docs/releases/SHOT_CHARTS_TIER0_COMPLETION_2026-07-31.md (repo docs, not public) methodology — Shot charts and spatial analytics — parked on permission, not failure ↗
Result RETAINED BASELINE
MLB champion model tournament — nothing promoted #
Rinehart Ratings · MLB · Correction 2026-08-08
Question Does any candidate MLB team-strength model earn promotion over simple baselines?
What happened
“No contestant separated from the field. elo took 3 of 4 outer seasons, and its overall margin_mae (3.4241) differs from champion-ridge (3.4229) by 0.0013 runs, inside the preregistered materiality floor of 0.01. On this evidence the models are not distinguishable, so nothing is promoted and no ranking between them is claimed.”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters Corrected evidence, 2026-08-08: the earlier bootstrap-CI framing is withdrawn — "these are sensitivity ranges—not 95% confidence intervals or tests of separation." The correction enforced strict completion-date availability and exactly-equal contestant masks over 8,841 common games. mlb-champion-ridge-v0.1.0 stays research only, never a rank claim.
Technical record
SRC Rinehart Ratings docs/audits/2026-08-08_MLB_CHAMPION_TOURNAMENT_CORRECTED.json (repo doc, not public; supersedes the 2026-08-06 artifact, which remains as historical evidence of the superseded method and claim) methodology — MLB champion model tournament — nothing promoted ↗
Result BLOCKED BY RIGHTS
MLB live data source — blocked until someone says yes #
Rinehart Ratings · MLB · Disposition 2026-08-06
Question Is there an approved live MLB data source for a 2026 Current Board?
What happened
“Status: BLOCKED. No live MLB source is approved, and none has been requested from a
provider.”
Decision Stop until a lawful source is admitted.
Why it matters The MLB 2026 Current Board is UNAVAILABLE by construction; the historical/research vertical remains active. "The honest answer up front: no candidate has clean written permission for this use." Clarified 2026-08-08: "“No live source exists” should be read as “no live source has been reviewed and approved.”" An MLBAM permission inquiry is drafted and unsent.
Technical record
SRC Rinehart Ratings docs/mlb-live-source-decision.md + docs/mlb-live-source-review-2026-08-06.md (repo docs, not public) methodology — MLB live data source — blocked until someone says yes ↗
Result RESEARCH RECORDED NULL
CORE-WAR accounting/scale receipt — a 12-hour NULL #
Court Signal · NBA · Disposition 2026-07-27
Question Does the CORE-WAR accounting and scale validation earn promotion?
What happened
“CORE-WAR receipt assembled — verdict NULL, 10/13 criteria, no promotion.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters Receipt 2b4c4d02… (CORE WAR ACCOUNTING SCALE VALIDATION, 10/13 criteria) is listed publicly in evidence.json. This entry previously conflated two receipts; the conditional-prediction receipt (2e8b9131…, 15/16) now has its own entry below.
Technical record
SRC Court Signal DECISIONS.md 2026-07-27 ("CORE-WAR accounting/scale pilot: NULL, no promotion"), receipt 2b4c4d02… (repo doc, not public; receipt listed publicly in courtsignal.rodericrinehart.com/data/evidence.json ) evidence — CORE-WAR accounting/scale receipt — a 12-hour NULL ↗
Result RESEARCH RECORDED NULL
Attribution experiment receipt — reproducible, and NULL #
Court Signal · NBA · Disposition 2026-07-28
Question Can possession-level attribution beat the frozen criteria it preregistered?
What happened
“Attribution receipt assembled — verdict NULL, 6/10 criteria, no promotion.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters The final experiment reproduced all 11 scientific outputs byte-for-byte before returning its NULL — reproducibility and promotion are separate bars, and only one was cleared.
Technical record
SRC Court Signal ROADMAP.md header block, receipt 62df00c4… (repo doc, not public) evidence — Attribution experiment receipt — reproducible, and NULL ↗
Result RESEARCH RECORDED NULL
BUCKETS formula pilot — first result-complete validation, NULL #
Court Signal · NBA · Disposition 2026-07-24
Question Does the BUCKETS formula earn a production nomination under frozen criteria?
What happened
“the BUCKETS pilot EXECUTED: first result-complete formula validation; verdict NULL”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters Terminal NULL with no production nomination. The product’s public evidence file states: "every completed pilot to date returned NULL."
Technical record
Result RESEARCH RECORDED NULL
CHEF formula pilot — second result-complete validation, NULL #
Court Signal · NBA · Disposition 2026-07-25
Question Does the CHEF formula earn a production nomination under frozen criteria?
What happened
“The CHEF pilot EXECUTED: second result-complete formula validation; verdict NULL”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters Terminal NULL. CHEF was also relabeled "Three-Point Production" after a 2026-08-03 integrity correction found the prior label misdescribed the formula.
Technical record
Result RESEARCH RECORDED NULL
MAMBA formula pilot — third result-complete validation, NULL #
Court Signal · NBA · Disposition 2026-07-25
Question Does the MAMBA (Creation Load) formula earn a production nomination?
What happened
“The KOBE pilot EXECUTED: third result-complete formula validation; verdict NULL”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters The program ran as kobe-formula-validation; its receipt displays publicly as MAMBA after the 2026-07-30 metric-identity migration (MAMBA — Creation Load). Terminal NULL, no nomination.
Technical record
SRC Court Signal DECISIONS.md 2026-07-25 heading (repo doc, not public); receipt a72ff69b… listed publicly (display label MAMBA) in courtsignal.rodericrinehart.com/data/evidence.json evidence — MAMBA formula pilot — third result-complete validation, NULL ↗
Result RESEARCH RECORDED NULL
ALIEN pilots — primary and DBPM-free comparator, both NULL #
Court Signal · NBA · Disposition 2026-07-31
Question Does an independent defensive component justify changing ALIEN — or removing its DBPM dependency?
What happened
“ALIEN primary/comparator complete — both exact replay, both NULL.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters "Removing DBPM materially weakens ALIEN. Retain the dependency and disclose it; no production formula change is nominated." The retained limitation — a disclosed DBPM dependency — is itself the finding.
Technical record
SRC Court Signal ROADMAP.md (repo doc, not public); receipts 3772bbf9… (primary, 13/19) and 36cdc277… (comparator, 7/14) listed publicly in courtsignal.rodericrinehart.com/data/evidence.json evidence — ALIEN pilots — primary and DBPM-free comparator, both NULL ↗
Result RESEARCH RECORDED NULL
Conditional CORE-WAR prediction — exact replay, 15/16, NULL #
Court Signal · NBA · Disposition 2026-07-29
Question Does conditional CORE-WAR prediction earn promotion at a frozen, matched task?
What happened
“Conditional CORE-WAR prediction complete — exact replay, 15/16, NULL.”
Decision Record the NULL. Do not treat no-change as a hidden success.
Why it matters "Full-precision CORE materially beats the training-mean null for margin and win probability; it is practically indistinguishable from the frozen two-decimal ablation." 2,767 games / 5,534 team-game rows with 5,000 paired whole-game bootstrap draws — distinct from the accounting/scale receipt above.
Technical record
Result FAILED
CFP committee rankability — the ordering did not validate #
Rinehart Ratings · CFB · Disposition 2026-08-08
Question Can a resume-based model reproduce the CFP committee’s final top-12 well enough to build a path model on?
What happened
“9.75/12 — FAIL”
Decision Keep the failure on the record. Do not publish this approach as a success.
Why it matters Preregistered bar: mean final top-12 overlap of at least 10.0/12. The nuance is real — a censored Plackett–Luce on generic resumes beats Elo-only and record-then-Elo in 11 of 12 held-out seasons, and still misses more than two field slots per season; in-season committee persistence beats every resume model.
Technical record
SRC Rinehart Ratings docs/research/2026_CFP_PATH_RESEARCH_II_DECISION.md + docs/research/2026_CFP_COMMITTEE_RANKABILITY.json (repo docs, not public) methodology — CFP committee rankability — the ordering did not validate ↗
Result FAILED
ACC/Miami partial identification — the bounds are not narrow #
Rinehart Ratings · CFB · Disposition 2026-08-08
Question Can tiebreak ambiguity be bounded tightly enough to publish a target-team path probability?
What happened
“0.208 / 0.180 — FAIL, under every policy reading”
Decision Keep the failure on the record. Do not publish this approach as a success.
Why it matters Miami’s title-game probability is identified only to [0.631, 0.839] against a preregistered width bar of 0.10. "An ACC clarification would narrow the bounds and still not rescue them." 44.9% of simulated seasons reach a tiebreak rung only SportSource can observe.
Technical record
SRC Rinehart Ratings docs/research/2026_CFP_PATH_RESEARCH_II_DECISION.md + docs/research/2026_ACC_MIAMI_PARTIAL_IDENTIFICATION.json (repo docs, not public) methodology — ACC/Miami partial identification — the bounds are not narrow ↗
Result RETAINED BASELINE
CFP Path research — stopped by its own stopping rule #
Rinehart Ratings · CFB · Disposition 2026-08-08
Question Is a CFP Path model justified on current evidence?
What happened
“STOP. Research III is not justified on current evidence.”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters Both preregistered decisive experiments failed; either alone would have ended the campaign. Research I remains fail-closed before full implementation with the owner hypothesis UNRESOLVED — "the current evidence does not authorize a target-team probability." No forecast snapshot exists and no public surface was deployed.
Technical record
SRC Rinehart Ratings docs/research/2026_CFP_PATH_RESEARCH_II_DECISION.md + docs/research/2026_CFP_PATH_MODEL_DECISION.md (repo docs, not public) methodology — CFP Path research — stopped by its own stopping rule ↗
Result RETAINED BASELINE
PULSE forward test — leakage found, evidence corrected, finding retained #
Hardball Signal · MLB · Correction 2026-08-08
Question Did the forward test behind PULSE’s anti-forecast lede hold up under audit?
What happened
“PULSE trailing-50 and trailing-100 were worse than the sealed league reference, season-to-date
was not distinguishable, and no EWMA survived selection-aware promotion”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters PM-01 (critical) found the v1 test estimated its league comparator from post-cutoff outcomes; two re-attacks found further leakage. The v2 evidence seals a target-independent cutoff reference, and the product’s live lede now reads: "The league reference is now frozen entirely before the target period, and paired intervals find the short windows have higher error. Season-to-date is not distinguishable from that reference."
Technical record
SRC Hardball Signal docs/releases/FINAL_REPORT_2026-08-08.md + docs/audits/POST_MOONSHOT_AUDIT_2026-08-08.md PM-01 (repo docs, not public) methodology — PULSE forward test — leakage found, evidence corrected, finding retained ↗
Result RETAINED BASELINE
Baserunning shrinkage tournament v2 — the winning arm was still not promoted #
Hardball Signal · MLB · Disposition 2026-08-08
Question Does partial pooling at K=25 earn promotion over the unpooled baserunning baseline?
What happened
“`K=25` beat unpooled but ranked behind league-only and its league interval spanned zero, so the arm was not promoted; v1 realized values remain unchanged”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters v1’s selection had scored aggregate full-sample means with no real split and no interval (PM-16); v2 reran a nested chronological evaluation before declining to promote.
Technical record
SRC Hardball Signal docs/audits/POST_MOONSHOT_AUDIT_2026-08-08.md PM-16 disposition (repo doc, not public) validation — Baserunning shrinkage tournament v2 — the winning arm was still not promoted ↗
Result RETAINED BASELINE
season_stats corrections — a withdrawn claim and an impossible display #
Hardball Signal · MLB · Correction 2026-08-08
Question What happens when a published claim or display is found wrong after release?
What happened
“The event policy produces 3,439 SB and 981 CS
versus official gamelog credits of 3,440 and 989 across exactly nine clean games.”
Decision Keep the incumbent. Do not promote the challenger.
Why it matters The exact official-gamelog conservation claim was withdrawn — the data was not (D-014); the producer chose versioned disclosure over reclassifying events to force agreement. The same successor wave removed the numeric ip field after impossible displays like "187.7" innings for 563 outs, which render as "187.2" in baseball notation (D-015). Every correction ships as a versioned successor artifact with predecessors preserved.
Technical record