Skip navigation
Dobitka

Does this model actually get it right?

A fair question, and you deserve it answered before you trust your own result. The model does not pick one outcome - it works out how often each one comes up. Every neutral base forecast is logged before kick-off and compared with what really happened. Everything below is the running tally of that benchmark, with no cherry-picking - not an evaluation of a specific user's plan.

How often the model is right

The page recalculates its metrics from the settled predictions available in the database on every visit.

29 neutral base predictions settled, from 28/08/2026 to 30/08/2026. The window covers up to the 3,000 most recent predictions: only model v0.4.3, season 2026/27 and supported club competitions. Every prediction was logged before kickoff.

Building the sample: 29 of 150 settled matches. The per-100 hit rates go public once the sample passes 150 matches - below that a single weekend can move the figure by double digits, so it would be noise, not a measurement. The raw metrics live in the section for the curious below, and the counter grows with every settled match.

29 / 150

For the curious: Brier, RPS and calibration

Hit rate alone is a blunt measure - it rewards safe favourites and cannot see whether the probabilities themselves are honest. These metrics check that.

Brier / RPS

0.639 / 0.229

It is not only whether the call was right, but how much confidence went into it. Two models back the home side and both are right - but one gave them 85%, the other 40%. That is the difference this shows, lower = better. Simple reference point on the same set: 0.634 / 0.225.

Raw-probability calibration

0.180

Whether "60%" really means 60%. We take every match forecast at 60% and check how many ended that way. 60 out of 100 - promise kept. Mean gap: 0.180 (closer to zero = more honest).

How we measure it

This sample contains only raw, neutral base predictions from the current model for the club season and supported competitions. The benchmark uses the predicted home XI, default playing style and both teams' strength; it does not contain a specific user's decisions. A prediction must exist before kickoff. After the match we compare the full 1X2 vector with the real result; in cup ties settled after a draw, the winner counts. An additional calibration layer is not applied to the current season's 1X2 result, points or ranking.

We improve the model through separate, versioned releases. Experimental variants do not change points or the ranking. Every change to the result path requires tests, an explicit model version and a human decision - the engine never switches automatically.

Latest settled match in this sample: 30/08/2026.

What the model takes into account

  • -The strength of both teams and their player profiles - from current squads and availability. The opponent's XI is added when it can be known before kick-off.
  • -Your plan: lineup, formation, playing style and planned substitutions. Each of these shifts the result distribution - otherwise there would be nothing to play for.
  • -Each of the eleven players separately: their xG and xA (shot and pass quality), their form over their own last five matches, and the position they actually play. We count player form, not club form.
  • -We run the same match a second time without your decisions, on the default setup. That is how you know how much you moved yourself, and how much was just the club's form.
  • -Thousands of paths through the same match, which form the distribution - for example "68 out of 100 simulations".
  • -The displayed result is the most common exact score. Minutes, scorers and the timeline come from a representative iteration that ended with that score.

What we do not count: weather, kilometres on the coach, derby temperature. The engine can do all of it, but those corrections are copied from the literature and we have not yet checked on our own settled matches whether they help at all. Until we do, they stay off. We would rather run a narrower model than a longer list.

The exact weights, the order of adjustments or the full feature list. That is the only thing we keep under the hood - and the only reason your result cannot be reproduced with a pen and a calculator.

See the model in action

Set a lineup and tactics, and the simulation shows whether your plan would have worked.

Is this page clear?