A team wins 1–0 and records 0.7 xG. Its opponent records 1.5. That is a useful reason to watch the match more closely. It is not permission to write "the winner was lucky" and stop thinking.
Perhaps the opponent generated much of its xG after falling behind. Perhaps a penalty explains the gap. Perhaps one provider sees a chance quite differently from another. The scoreline and the chance model are answering different questions.
An xG audit should help identify those questions. It cannot serve as a lie detector for a team's character, a manager's skill or the honesty of a prediction service.
First, identify the measurement
Expected goals assigns a scoring probability to a recorded shot. Features commonly include location, angle, body part and the preceding action. Some models also include goalkeeper and defender positions.
Hudl StatsBomb's explanation gives a helpful example: a basic model might rate a shot at 0.30, while a richer model that knows the goalkeeper is out of position could rate it at 0.65. Those are the provider's explanatory examples, not a finding that one provider is always right.
When totals differ, do not average the numbers to manufacture agreement. Establish whether the providers counted the same shots and used the same definitions. If an analysis switches provider halfway through a trend, disclose the switch.
xG also excludes dangerous moves that never became shots. A missed cutback and a successful tackle before the shot may tell a tactically important story without adding to conventional shot-based xG.
Use a small audit table, not a verdict
The following is a fictional five-match example. The numbers illustrate interpretation; they are not a club's results or an XG Mind performance claim.
| Observation | Team A | Team B |
|---|---|---|
| Goals scored | 10 | 4 |
| xG for | 6.0 | 7.0 |
| xG against | 7.0 | 5.5 |
| xG difference | −1.0 | +1.5 |
A scoreboard-only reading would favour Team A's attack. The xG totals raise a different question: how did Team A turn those opportunities into ten goals, and why did Team B convert so few?
Possible explanations include finishing quality, goalkeeping, event-model limitations and ordinary variation. These totals cannot separate the explanations on their own. They also cannot tell us whether Team A's next opponent is well suited to exploit its defence.
Treat the table as a prompt to investigate, not as a forecast that Team B must immediately improve.
Regression is not a deadline
Regression to the mean is often described as a debt: a team has scored too many goals and must now "pay them back." That is not how probability works.
An unusually strong observed stretch may contain both a genuine signal and positive noise. When the observation was selected precisely because it was extreme, later observations may be less extreme. That does not mean the next match reverses the previous result.
There is no universal six-match threshold, no rule that 135% of xG predicts an imminent collapse, and no guarantee that a team's goals-to-xG ratio converges to exactly one. A persistent finisher, a model that omits a relevant feature, or a change in personnel can complicate the comparison.
Use a wider window, inspect the shot mix and state what the evidence cannot distinguish. A restrained sentence such as "the scoring run is stronger than this provider's recorded chance quality" is more informative than "the crash is coming."
Three checks before interpreting a difference
Penalties. Keep total xG and non-penalty xG separate when appropriate. Do not mechanically subtract a universal constant. StatsBomb documents that penalty values can differ between models and versions. Remove the actual penalty-shot contribution recorded by the chosen source. Winning penalties can also reflect attacking behaviour; they are not automatically "unearned."
Game state. A team protecting a lead may allow more territory and shots. Inspect when the chances happened, red cards and score changes. A post-match total can describe the game that unfolded without representing the pre-match strength balance.
Opposition and sample. Five fixtures against weak opponents are not interchangeable with five against strong opponents. The choice of competition, home/away split and time window matters. Changes in coach or lineup may make an older sample less relevant, but a very short newer sample brings its own uncertainty.
These are checks, not automatic adjustments. If you claim to have corrected for game state or opposition, explain the procedure rather than simply placing "adjusted" in front of the metric.
What post-shot xG adds
Post-shot models incorporate information about the shot after it has been struck, such as placement and sometimes velocity. They are often used to study goalkeeping. The names PSxG and xGOT do not guarantee identical definitions across providers.
Under a compatible convention, post-shot xG faced minus goals conceded can describe a keeper's performance relative to that model over a sample. A positive value is not proof of permanent skill; a negative value is not proof of a poor keeper. Definitions, excluded events and sample size still matter.
One spectacular match is worth describing as one spectacular match. It does not establish an inevitable reversal in the next one.
Keep the forecast and the audit separate
An audit after full-time has information unavailable before kickoff. It can explain recorded chance creation, but it cannot retrospectively validate a forecast unless the original forecast and its inputs were preserved.
For readers, the practical routine is short: identify the provider, inspect penalties and game state, check the fixture window, and read the shot pattern. Then ask whether the written conclusion is stronger than the evidence.
Our team pages offer a place to inspect the available data. The missing-data note covers what to do when a comparison lacks one side's metrics.
Source and correction
- Hudl StatsBomb: expected goals explained, including provider differences, penalties and post-shot xG.
Correction, 1 October 2026: An earlier version gave unsourced correlation coefficients, a Union Berlin case study without a verifiable dataset, and fixed thresholds for predicting regression. Those claims have been removed. This article now uses an explicitly fictional example and retains its original URL.