An analyst asks for a club's average xG over five matches. The database has values for three. Two fields are empty.
One query replaces the empty fields with zero and divides the total by five. Another silently averages the three available values and calls it a five-match average. Both outputs look tidy. Neither accurately describes the evidence.
Missing-data handling is not housekeeping at the edge of football analysis. It can change the central conclusion while leaving no obvious trace in the finished preview.
Put the denominator beside the number
Take a fictional sequence of xG observations: 1.4, missing, 1.8, missing, 1.0.
The observed values sum to 4.2. Dividing by five gives 0.84; dividing by the three observed fixtures gives 1.40. The second is a valid mean of the observed values, but it is not automatically an estimate of the full five-match period.
| Treatment | Output | What the reader should be told |
|---|---|---|
| Replace missing with zero | 0.84 | Two missing observations were assumed to be zero; usually inappropriate. |
| Mean of observed fixtures | 1.40 | Only three of the five fixtures have this metric. |
| Apply a documented fallback | Depends on the rule | Which values were estimated, how and from which evidence. |
| Omit the comparison | No average reported | Coverage is too limited for the intended claim. |
The missing fixtures may not resemble the observed ones. A provider might cover major league games but omit cup games. Complete-case averages can therefore change the competition mix rather than merely reduce the sample size.
Separate the different kinds of absence
An empty field can mean that the provider does not cover the competition, that the match's statistics have not arrived, or that our collection failed. It can also mean an identity-mapping problem: a row exists, but it belongs to a different team in the source's ID system.
Those situations require different responses. Retrying may resolve a delayed update. It will not create tracking metrics for a competition the source does not cover. Filling a mapping failure with an average conceals an engineering fault.
Stale is another state, not a synonym for missing. Yesterday's team totals may be valid for yesterday's cutoff but not for a report that claims to include today's completed fixture.
A practical record should distinguish observed, missing, estimated and stale values. Preserve a reason when it is known. Where the reason is unknown, say so instead of guessing.
Define completeness for the question being asked
A non-empty statistics table is not necessarily a complete statistics table. It may contain shots and possession while omitting xG. A "last five" sample may contain four settled matches and a postponed fixture.
Check the fields required by the actual analysis. If the claim concerns non-penalty xG, total xG alone is not enough. If a keeper comparison requires post-shot metrics, conventional xG cannot silently replace them.
This distinction prevents a common software error: treating "some rows exist" as proof that the requested dataset is ready. Coverage should be a property of named measurements over a specified period.
The xG audit guide explains why the provider and definition belong alongside those measurements.
Fallbacks change the question
Early in a season, last season's data may offer context when the current sample is small. It is a defensible choice only if it is disclosed. A promoted team, changed coach or heavily rebuilt squad can make the old sample poorly matched to the current one.
A league-level baseline is another possible fallback. It describes a prior expectation for the relevant population, not a newly observed fact about the club. If the team has little evidence, the report should not present the baseline to two decimal places as though it were a precise club measurement.
A blend of old and new observations also needs a rule. Record the weights, eligibility and evaluation procedure. Choosing the blend that produces the preferred narrative is not a method.
Sometimes the best fallback is to leave the comparison out. There is no requirement that every report contain every advanced statistic.
Show uncertainty without inventing confidence
A useful limitation statement is specific: "xG is available for three of the five league fixtures; the two missing matches are excluded from this comparison."
An unhelpful one says "data may be incomplete" after presenting a confident ranking. The reader cannot tell which part of the ranking is affected.
Reduced coverage may justify more cautious interpretation, but it does not mechanically imply a particular numerical confidence reduction. Unless a procedure for that reduction has been tested, avoid inventing an adjustment such as "subtract 10% confidence."
Language models need the same information. Give them the metric value, available count, requested count, period and missing reason where known. A fluent writer cannot recover observations that were never supplied.
Preserve the fallback used at the time
An upstream provider may fill the empty field a day later. Updating the current database is sensible. Reconstructing an old forecast from that improved database is a separate operation.
Keep the original input snapshot or at least a versioned record of the values and fallback used. Otherwise a later audit can make yesterday's system appear to have known more than it did. Our data-leakage article examines that timing problem in detail.
If a missing value affected a published conclusion, make the correction visible. Fixing the input without acknowledging its effect can leave the reader with an unexplained reversal.
What to display in a data-driven preview
For an important rolling metric, show the provider, competition, period, observed fixture count and last relevant update. Mark fallback values distinctly. Distinguish "not covered" from "temporarily unavailable" where possible.
These are design recommendations, not a claim that every XG Mind data page already displays all of them. The purpose is to give readers a standard against which any platform, including ours, can be checked.
Further reading: Scikit-learn: imputation of missing values describes general imputation approaches. Choosing one for football still requires domain assumptions and out-of-sample evaluation; a software default is not a justification.