An 80% hit rate might describe eight correct calls from ten easy fixtures. It might describe a much larger, genuinely difficult test. It might also hide excluded draws or unpublished failures. The number alone does not tell us which.
There is a tempting shortcut: dismiss every claim above 65% as mathematically impossible. That shortcut is wrong. Football is uncertain, but uncertainty does not create one fixed accuracy limit across all competitions, fixture selections and prediction tasks.
This article corrects our earlier claim of such a limit. The better way to challenge a performance claim is to ask for its denominator and compare it with a baseline on exactly the same matches.
Start with the task, not the percentage
A single home/draw/away choice is a different task from choosing two of those outcomes. Predicting an exact score is different again. A measure that counts any one of several suggestions as a success will not be comparable with a single-choice measure.
Selection also changes the test. A method that only calls strong favourites is allowed to achieve a high hit rate; it simply has not been evaluated on a representative full schedule. That matters if the headline suggests general football forecasting ability.
| Question | Why it changes the interpretation |
|---|---|
| Were all fixtures included? | Excluding difficult matches changes the population. |
| Was each call published before kickoff? | Later information can improve an apparent historical record. |
| Were draws counted? | Removing a class changes both the denominator and the task. |
| Did a simple baseline do just as well? | High accuracy need not represent an improvement. |
| Are the selection rules fixed? | Rules chosen after seeing results introduce bias. |
A percentage can be numerically correct and still communicate the wrong impression.
A Poisson model does not stop at 65%
For a basic illustration, suppose home and away goals are independent Poisson variables with mean rates lambda_home and lambda_away. A Poisson variable has probability:
P(G = k) = exp(-lambda) × lambda^k / k!
The home-win probability is the sum of the joint probabilities for score pairs where home goals exceed away goals. There is no constant 0.65 in this calculation.
The following values are constructed examples, not estimates for named clubs or findings from an XG Mind backtest:
| Home goal rate | Away goal rate | Calculated home-win probability |
|---|---|---|
| 1.4 | 1.2 | 41.56% |
| 2.0 | 0.8 | 65.44% |
| 3.0 | 0.3 | 90.24% |
The last row is already a counterexample to the proposed mathematical ceiling within that model. It does not mean the rate estimates describe any particular real fixture accurately. It means the distribution itself does not impose the claimed limit.
The figures can be reproduced using the Python standard library:
from math import exp, factorial
def poisson_mass(goals, rate):
return exp(-rate) * rate**goals / factorial(goals)
def home_win_probability(home_rate, away_rate, max_goals=40):
return sum(
poisson_mass(home, home_rate) * poisson_mass(away, away_rate)
for home in range(max_goals)
for away in range(home)
)
for rates in [(1.4, 1.2), (2.0, 0.8), (3.0, 0.3)]:
print(rates, round(100 * home_win_probability(*rates), 2))
This uses scores from zero to 39; the omitted tail is negligible for these particular rates. Independence and constant goal rates are simplifying assumptions. Real models may account for dependence, game state and other effects. None of that rescues the universal ceiling claim.
What the relevant accuracy limit would depend on
Imagine knowing the true probabilities of home win, draw and away win for every fixture. Under ordinary single-choice accuracy, the best choice for each match would be the outcome with the highest probability. Its expected success rate would be the average of those highest probabilities over the selected fixtures.
Change the fixture population and that average changes. This is why balanced fixtures and extreme mismatches cannot share an assumed universal limit.
In practice we do not know those true probabilities. We estimate them, and those estimates can be wrong. A defensible evaluation therefore needs held-out results and a clear baseline rather than a claim that the model has reached "football's mathematical frontier."
High accuracy can coexist with poor probabilities
Two systems may always choose the same winner while assigning very different probabilities. One says 60%, the other 95%. Their hit rates are identical, but their confidence is not.
If the chosen outcome happens around 60% of the time, the 95% system is badly overconfident. A reader who sees only the final direction will miss this weakness. Our calibration guide gives a worked example using Brier scores and reliability checks.
Conversely, losing one match after publishing a 70% probability is not a contradiction. The forecast allowed for failure. Calibration asks whether comparable events occur at roughly the stated frequency, not whether every individual forecast succeeds.
A compounding argument is not a proof of impossibility
The earlier version also used a hypothetical financial compounding calculation to reject high hit rates. Its numerical result was incorrect, and the inference was too strong.
Accuracy alone does not determine a return. Prices, costs, selection, dependence, limits and estimation error all matter. Even a properly specified expected-growth calculation does not establish a guaranteed realised path. It cannot prove that a particular accuracy percentage is impossible.
For a research blog, the useful conclusion is simpler: don't replace an audit with a dramatic number. Ask for the forecast history, evaluation rules and uncertainty.
The XG Mind public record should be read with those questions too. A platform's own statistics deserve the same scrutiny as anyone else's.
Reading and correction
- Scikit-learn: probability calibration, including the distinction between calibration and broader predictive quality.
- Scikit-learn: time-series splits, on testing ordered observations without training on future data.
Correction, 1 October 2026: This replaces the unsupported "65% mathematical ceiling" argument and removes the erroneous compounding calculation. The example probabilities above were calculated directly from the stated assumptions. The URL is unchanged.