In 2026, the picks ran to a longest streak of 7.

Across 101 scored days it picked 98 and skipped 3, hitting on 70.4% of picks, with confidence calibrated to a Brier score of 0.2378 (lower is better; 0.25 is a coin flip).

Walk-forward backtest — every day scored by a model trained only on days before it, then the real skip / pick / Double Down policy run day by day against a running streak.

Season

Streak length, day by day — 2026

Climbs +1 per hit, drops to zero on a miss (a Double Down miss included), holds flat on a skip.

0367May 14Jul 3Aug 26

Confidence vs. reality

Each dot is a confidence bucket, sized by how many predictions it holds. On the dashed line, the percentage told the truth. Buckets with fewer than 30 predictions are left off as too thin to read.

40405050606070708080PREDICTED %

Every season

SeasonBrierLongestMeanHit ratePicked
20230.2365113.175.9%58
20240.2380143.169.4%157
20250.2369173.470.5%156
20260.237873.170.4%98

Last 30 scored days — 2026

hit miss skipped

Walk-forward backtest: each day scored by a model trained only on prior days. These predictions come from the baseline logistic-regression family (see CLAUDE.md); the deployed pick model is the three-way ensemble. Streak simulation runs the real decision policy day by day.

What "walk-forward" means

Every day in the 2026 column was scored by a model trained only on days before it — never on its own future. That's the honest way to ask "would this have worked," and it's why the numbers here are lower than a model graded on data it already saw. The streak figures come from running the real decision policy — skip, pick, and Double Down — day by day against a running streak.