How the Oracle predicts a cricket match
Short version: it rates every team from every completed match it can see, accounts for who is at home, blends in recent form and head-to-head, and turns that into a probability. It was judged on matches it never saw during training — and every live call is locked before the start so the record can't be edited after the fact.
1 · The data
Every international and franchise match in the Cricsheet ball-by-ball archive — over 15,000 men's and women's Tests, ODIs, T20Is and T20 league games from 2002 to today — plus results we record ourselves between Cricsheet updates. Fixtures come from ESPN's live scoreboard.
2 · Team strength — an Elo rating
Each side has a rating that moves after every match by how surprising the result was: an upset moves it a lot, an expected win barely at all. Ratings are kept separately for each format and for men and women — a men's side never borrows its women's results. Bigger winning margins count for more. Home advantage is measured, not assumed: in our data hosts win 64% of decisive men's Tests, 59% of ODIs and 53% of T20Is, and the model learns how much that is worth in each format.
Associate nations start below the twelve ICC Full Members. Without that, a run of wins in a regional qualifier would lift a side above established nations it has never beaten — the backtest confirmed the prior makes predictions measurably better.
3 · The prediction
The rating gap (including home advantage) is combined with recent form and head-to-head by a model fitted separately for Tests, ODIs, international T20s and franchise T20s — these behave very differently (franchise squads reshuffle every auction). Where a segment has too little history, the model uses the calibrated rating alone rather than risk learning noise. Tests also get a draw probability: evenly matched sides draw more often.
4 · How it was tested
The model was built and tuned only on matches from 2019–2023, then scored once on every match from 2024 onwards — 4,159 games it had never seen. Each prediction used only information available before that match. Only decisive matches between sides with at least 10 prior games were scored.
| FORMAT | MATCHES | PREVIOUS | ORACLE V2 | BRIER |
|---|---|---|---|---|
| Tests | 99 | 64.7% | 67.7% | 0.216 |
| ODIs | 435 | 62.5% | 66.0% | 0.206 |
| T20s (internationals + leagues) | 3,625 | 64.4% | 66.7% | 0.206 |
| All | 4,159 | 64.2% | 66.7% | 0.207 |
Accuracy is the share of matches where the favourite won. The Brier score measures whether the probabilities themselves are honest — a model that says 70% should be right about 70% of the time. It is the number we optimise, because a confident wrong call should cost more than a cautious one.
5 · The locked-call guarantee
Every prediction is locked 48 hours before the scheduled start. The rule is enforced by the database itself: it rejects any prediction created after a match starts and any edit to a prediction once made — only the result can be recorded afterwards. Before the lock you'll see a provisional view that updates as new results arrive; provisional views are never counted. The live record includes every locked call, right or wrong.
6 · What it can't do
- It doesn't know the playing XI, injuries, pitch or weather — only results. A rested star or a raging turner can beat it.
- Franchise T20 is hard for any model (~55% in our tests): squads change every season and matches are short.
- No call when the data is thin — sides need 10+ matches in the format. Afghanistan isn't in the open Cricsheet archive, so we don't predict their matches yet.
- Domestic first-class and List A cricket isn't covered.
For entertainment and analysis only — CricMind predictions must not be used for betting or wagering. Model file and validation are regenerated on every retrain; last trained on data through 2026-10-07.