Research & Engineering • Published: October 01, 2026 • Updated: Oct 02, 2026 • 5 min read

Same Winners, Different Quality: Checking Forecast Calibration

XG Mind Editorial
XG Mind Editorial Website →
Football data and research notes
Key Takeaways / Abstract

Two football models can pick the same winners and still differ sharply in quality. A worked Brier-score example shows how to check confidence, not just accuracy.

Two previews both predict a home win. One gives it a 60% chance, the other 95%. The home side wins. Both get a tick in the results table.

That table has lost an important piece of information. The second preview was almost certain; the first left considerable room for a draw or defeat. Over many matches, that difference becomes testable.

A direction hit rate measures whether the selected outcome happened. Calibration asks whether the confidence was warranted. If a website publishes probabilities, readers should be able to examine both.

Two different-sized lenses inspect a football pitch above a balanced scale, a metaphor for checking confidence against evidence.
Confidence needs to be checked against observations. AI-generated conceptual illustration, not a calibration chart; the numerical example follows below.

A small example with all assumptions visible

Take 100 fictional binary forecasts of "home win" versus "not home win." Exactly 60 home wins occur. Model A assigns 0.60 to every home win; Model B assigns 0.95. Both always select home win, so both have 60% direction accuracy.

For this binary event, the Brier score is the mean of (p − y)^2, where y is 1 for a home win and 0 otherwise. Lower is better.

Model Forecast probability Home wins observed Direction accuracy Binary Brier score
A 0.60 60 of 100 60% 0.2400
B 0.95 60 of 100 60% 0.3625

For A, the calculation is 0.60 × (0.60 − 1)^2 + 0.40 × 0.60^2 = 0.24. For B, it is 0.60 × (0.95 − 1)^2 + 0.40 × 0.95^2 = 0.3625.

A's probabilities match the observed frequency in this constructed sample. B is overconfident. A winner-only table cannot show that difference. Neither model in this example distinguishes easier fixtures from harder ones: every probability is the same.

That last detail matters. A constant forecast can match the overall frequency while being unhelpful for ranking individual matches.

Use the right version of the score

Football match direction normally has three classes: home win, draw and away win. A full probability forecast must assign each a non-negative probability, with the three adding to one.

One multiclass Brier convention averages the sum of the squared errors across those three classes. Its range is zero to two. Other reporting conventions rescale the score, so state the convention before comparing numbers.

Do not take the binary table above and present it as a three-way result. It deliberately merges draws and away wins into "not home win" to keep the arithmetic easy to inspect.

Log loss is another proper scoring rule. It penalises assigning very little probability to an outcome that actually occurs. If clipping is used to avoid numerical problems at zero, disclose it. Changing the clipping rule can change a reported score.

A reliability diagram needs counts

To inspect calibration, group predictions into probability ranges and compare each group's mean predicted probability with its observed event rate. A group averaging 0.60 should have an observed rate near 0.60, subject to uncertainty.

Publish the number of forecasts in each group. A point based on eight matches should not look as conclusive as one based on eight hundred. Too many narrow bins create noisy patterns; broad bins can conceal important variation.

For a three-way football model, inspect the classes separately or use an explicitly described multiclass approach. Home-win calibration alone does not reveal how well the model handles draws.

The scikit-learn calibration guide explains reliability diagrams and adds an important caution: a lower Brier loss does not necessarily mean better calibration. It also reflects discrimination and the underlying uncertainty of the task.

Compare like with like

A fair comparison uses the same held-out fixtures, publication horizon and result definition. Compare a twelve-hour forecast with a baseline that had access to twelve-hour information, not with one updated after the team sheets.

Useful baselines might include a training-period class frequency or a simple model fixed before evaluation. Whatever baseline is chosen, record the selection procedure. Do not choose the easiest comparison after seeing which model won.

Look at the full sample and sensible subgroups. A pooled reliability plot can hide overconfidence in one league and underconfidence in another. Conversely, slicing repeatedly until an attractive pattern appears can create its own misleading result.

If neighbouring fixtures share teams or errors, uncertainty estimates that assume complete independence can be too optimistic. State the resampling or interval method rather than labelling every difference "significant."

Calibration cannot be invented after the match

Suppose an archived preview contains only "home win" and a high-confidence badge. It is tempting to turn that badge into 80% after the event and calculate a Brier score.

That creates a new forecast that was never published. Unless the numerical mapping was specified and recorded beforehand, it is not a valid probability-performance record.

For XG Mind, our public results page should therefore not be called a calibration dashboard merely because it displays direction and score outcomes. Probability metrics require the relevant pre-match probability vectors to exist and be preserved. This article is a guide to evaluation, not a claim that the site currently exposes every metric discussed here.

A short review routine

Before accepting a model-comparison chart, check five things: the task, the held-out sample, the score convention, the forecast horizon and the count behind each calibration point. Then read the errors rather than only the headline winner.

If a model looks unusually accurate, the data-leakage guide is the next stop. Better-looking probabilities are not necessarily better forecasts if tomorrow's information entered yesterday's inputs.

Source: Scikit-learn: probability calibration. All numerical observations in this article come from the labelled fictional example, not a live XG Mind backtest.

Frequently Asked Questions

What is probability calibration in football forecasting?
Calibration compares stated probabilities with observed frequencies across comparable forecasts. Among many events assigned about a 60% chance, approximately 60% should occur, allowing for sampling uncertainty.
Does a lower Brier score prove better calibration?
Not by itself. Brier score reflects calibration, discrimination and outcome uncertainty. Compare scores on the same held-out matches and examine reliability diagrams and sample sizes as well.
Can direction-only predictions be used to calculate a Brier score?
Not without imposing new assumptions. A Brier score requires recorded probabilities. A direction label or a low/medium/high confidence label is not automatically a numerical probability distribution.