Model: pooled-poisson-v3 | Backtest: 1,183 matches, no lookahead | Live: 143 priced matches | Date: 2026-09 | Status: signals in shadow mode The model won its backtest. It beat league base rates, beat its predecessor, held up out of sample. Then it went live, fired 27 signal bets, and lost 6.2 units. The one benchmark I had not evaluated it against was the bookmaker — and the bookmaker won.
What I built
A small dashboard for Czech Liga and FNL football. One time-weighted attack/defence rating per team, pooled across both leagues with a one-year half-life, so promoted and relegated teams carry their rating between divisions. Ratings are shrunk toward the league average with pseudo-matches, a scoreline probability matrix comes out the other end, and 1X2 / Over-Under / BTTS probabilities fall out of the matrix.
Backtest, 1,183 matches, 2024–2026, no lookahead. Log loss, lower is better:
| Model | 1X2 | OU 2.5 | BTTS |
| League base rates | 1.072 | 0.694 | 0.694 |
| home-away-poisson-v2 | 1.024 | 0.692 | 0.695 |
| pooled-poisson-v3 | 1.013 | 0.689 | 0.692 |
That looks like progress. v3 cuts 1X2 log loss by 5% over base rates and beats v2 everywhere. If your mental model of "is the model good" is this table, you ship it.
The baseline I skipped
Bookmaker odds are probabilities with a margin stapled on. Invert them, normalize, and you have the market's own forecast — for free, no ratings needed:
def devig(odds):
# implied probs include the bookmaker margin, so they sum to > 1
inv = [1 / o for o in odds]
margin = sum(inv) # typically ~1.05-1.08 on 1X2
return [p / margin for p in inv] The de-vigged consensus of available books is the single strongest baseline that exists for a sport like football. Millions of stake flow through those prices; the margin is the tax, the normalized probabilities are the crowd's model. Typical 1X2 margins run 5-8%, which means "beat the raw odds" is a meaningless bar — you beat the odds by more than the margin or you have nothing.
I evaluated against base rates and my own older model. I did not evaluate against this. That is how a model ships while being worse than a four-line function.
143 matches later
Live, on priced matches with pre-kickoff snapshots only (historical odds are never backfilled — that's how you lie to yourself with backtests):
| Forecaster | 1X2 log loss, live |
| De-vigged bookmaker consensus | 1.006 |
| pooled-poisson-v3 | 1.009 |
| home-away-poisson-v2 | 1.025 |
v2 was losing to the market by 2 points of log loss and its signals were betting anyway: 27 bets, −6.2 units. The mechanism is simple once you see it. A signal fires when model probability × decimal odds > 1. If the model is overconfident by even 3 points on a selection, it manufactures expected value that does not exist. You don't have an edge. You have a bug that places bets.
v3 closes most of the gap (1.009 vs 1.006) and still does not beat the market.
The 10,429-bet version of the same lesson
A second engine — same sport, bigger scope, a season-long value-betting run — had already run this experiment the expensive way. Raw model, 10,429 backtested bets, −10.0% ROI.
The loss anatomy was the interesting part. Splitting predictions into deciles of model confidence and comparing against outcomes showed global overconfidence: +3 to +10 percentage points, in every decile. Not "bad at longshots" or "bad in one league." The whole calibration curve was shifted. Every EV computation built on those probabilities was inflated, so the engine preferred exactly the bets it was most wrong about.
Pre-registered fixes, applied and re-measured: cap odds at 2.50, drop away-win picks, drop January picks. Result: −6.73% over 2,942 bets, confidence interval −10.5 to −3.0. Still negative; the 2024+ half reached −1.1%. Better is not profitable.
One variant deserves its own line: a draws-only filter showed +3.2% ROI. Its confidence interval ran −15 to +21. Adopting a rule because a noisy slice of a losing strategy printed a positive number is how losing systems get more complicated. Rejected.
What runs now
Signals are rebuilt market-first, and they stay in shadow until the model earns control:
p = 0.7 * market_prob + 0.3 * model_prob # blend, market-anchored
signal = (
p * best_odds >= 1.03 # +3% EV at the blended prob
and abs(model_prob - market_prob) <= 0.07 # else: assume the model is wrong
) - 1X2 only. On OU 2.5 and BTTS the improvement over base rates (0.694 → 0.689) does not justify a bet.
- Disagreement over 7 points kills the signal — that is almost always model error, not value.
- Teams that changed division wait 8 matches in the new division before they can signal at all.
- Every candidate is snapshotted pre-kickoff and settled per model version, log loss and Brier against the same de-vigged consensus.
- Nothing goes live as a signal until the current model version beats the market's 1X2 log loss over 300 settled, priced matches. Until then: shadow.
What I found
- The market prior does the heavy lifting. A 70/30 blend toward the market removed the manufactured-EV failure mode on day one.
- The disagreement gate (>7 points) catches model errors before they cost anything — it fires on exactly the fixtures where v2 was most confident and most wrong.
- Both engines failed the same way: not accuracy, calibration. v3's forecast quality is real (1.013 vs 1.072 base rates) and commercially irrelevant next to 1.006.
- A positive slice with a CI through zero (+3.2%, CI −15 to +21) is noise wearing a sign that says EDGE.
What I'd change
De-vig the market on day one, before writing any model code — it is four lines and it is the bar everything else has to clear. I spent the effort order exactly backwards: ratings first, baselines second, market last. The shadow-mode rule is the other thing I'd keep forever: a model that has never beaten the market on settled matches has no business placing bets, no matter what the backtest says.