A useful forecast archive lets a sceptical reader ask more than "how many winners did you get?" It lets the reader reconstruct what was predicted, when it was available and how it was scored.
FiveThirtyEight's Soccer Power Index remains a useful reference for that idea. Its data documentation describes match-level ratings and forecasts dating back to 2016, including home-win, away-win and draw probabilities, projected goals and eventual scores.
The archive's historical value does not depend on calling it the first or the best football model. It comes from being inspectable. Nor should a preserved dataset be mistaken for a currently updated service: check its coverage and dates before using it as a live source.
A record needs a unit of evaluation
Start with the thing being scored. Was the forecast a single match direction, an exact score, a goals total, or a probability distribution? A record should not quietly change that definition when the result is inconvenient.
A minimum archive needs a stable fixture identity and a link between the published prediction and the settled outcome. The dataset should also state how it handles postponements, abandoned games, extra time and corrected official results.
| Record field | Question it helps answer |
|---|---|
| Fixture identifier | Are these entries referring to the same match? |
| Publication timestamp and version | What could a reader actually see before kickoff? |
| Prediction and evaluation rule | What counted as success at the time? |
| Outcome and source | Which result was used to settle the forecast? |
| Correction history | Was anything changed, and why? |
| Included and excluded entries | Does the headline describe the entire published set? |
A precise denominator is often more valuable than another decimal place in the hit rate.
Public is not the same as immutable
A web page can be publicly visible and still editable by its operator. An application-level edit restriction may prevent ordinary users from changing a settled report while leaving administrators or database operators able to change it.
That distinction is important for XG Mind too. Our results page exposes settled results, but visibility alone is not evidence of cryptographic immutability. We should not claim that ordinary relational rows are "write-once, read-forever" without demonstrating the mechanism.
Possible stronger controls include retaining version history, publishing forecast-file digests before kickoff, or keeping an independently timestamped copy. Each has limitations. A digest can help demonstrate that a file has not changed after the digest was published; it does not prove that its contents were accurate or that every forecast was included.
Corrections should remain possible. A wrongly assigned fixture, delayed official result or broken scoring rule needs a fix. Preserving the original version alongside a dated explanation is more honest than pretending that software never makes mistakes.
Losing streaks need arithmetic, not theatre
A losing run is not automatically proof of failure. It is also not proof that a model is genuine.
Here is a constructed probability example, not an observed XG Mind result. Suppose every prediction succeeds independently with probability 0.58. The failure probability is 0.42. The probability that one particular group of four calls all fails is:
0.42^4 = 0.03111696, or about 3.11%.
That is not the probability of a four-loss run appearing somewhere in 100 calls. There are overlapping windows, so treating all 97 windows as independent would be wrong. A small recurrence tracks the probability of reaching the next call without having completed a four-loss run:
def probability_of_failure_run(trials, failure_probability, run_length):
states = [1.0] + [0.0] * (run_length - 1)
for _ in range(trials):
following = [0.0] * run_length
following[0] = sum(states) * (1 - failure_probability)
for length in range(1, run_length):
following[length] = states[length - 1] * failure_probability
states = following
return 1 - sum(states)
print(round(probability_of_failure_run(100, 0.42, 4), 6))
The result is 0.853902, about 85.39%. Those assumptions matter: actual fixtures can have different probabilities and correlated errors. This calculation is not an estimate for a particular team's schedule or our own model.
The earlier version mixed incompatible figures in the FAQ and added longer-run percentages without a calculation. A numerical example should not acquire authority simply because it looks technical.
Report the full sample, then the useful slices
Readers may want performance by league, season or forecast horizon. Those slices can reveal weaknesses hidden by a pooled number. They can also encourage cherry-picking if the operator only displays favourable groups.
Show the full sample first, state the segmentation rule and include group sizes. Ten successful calls in one competition should not carry the same interpretive weight as a much larger history. If the method changed mid-season, record the version rather than blending everything into an uninterrupted claim of improvement.
This is also why pre-match publication matters more than a universal twelve-hour deadline. A forecast at twelve hours and one at twenty minutes are different products with different information available. Both can be evaluated honestly if their horizon and versions are recorded. Neither requires a claim of guaranteed settlement within thirty minutes of the whistle; upstream result feeds can be delayed.
What readers should be able to take away
A trustworthy archive makes disagreement easier. A reader should be able to calculate a baseline, inspect failures and challenge an exclusion rule without relying on the operator's sales copy.
That is the useful lesson to carry forward from inspectable football datasets. It does not require claiming to replace FiveThirtyEight, possess a multi-season record that does not yet exist, or reproduce another organisation's model.
For a closer look at the evaluation itself, read our probability-calibration guide. For the timing problem, the data-leakage guide explains why a current database is not necessarily a faithful archive of what was known before a match.
Source and correction
- FiveThirtyEight SPI data documentation, inspected through GitHub's public contents API because the directory page disallows automated fetching.
Correction, 1 October 2026: This article no longer claims that XG Mind has a technically immutable database, a guaranteed publication horizon or settlement deadline. It distinguishes transparency from immutability, reconciles the streak example and removes unsupported descriptions of other organisations.