Open methodology
How the arena is scored
Same fixture context for every model, one prediction each, locked before play, judged on the 90-minute result. The site publishes enough to let you recompute any number on any board, which is the point of writing this page down.
Three boards, and what each one answers
The site carries three leaderboards. They rank the same kind of thing and they do not always agree, which is intended: they are answering different questions on different samples. The grading rules below are shared by all three. What changes between them is which models and which fixtures go in.
| Board | Models | Fixtures | The question it answers |
|---|---|---|---|
| The study | 13, frozen | 1,000, fixed in advance | Do these models differ from each other, and from the market, on evidence nobody chose after the fact? |
| The full record | 13, the final field | Every fixture the site has graded | What does the field look like right now? |
| World Cup 2026 | 11 | 104 matches, complete | How did the models call one tournament, start to finish? |
If you only want one number, the study is the one to trust, because its sample was settled before any of it was played. It is also the slowest to fill.
What we ask a model for
We never ask a model which team will win. We ask two separate questions, and it answers both in one JSON object:
{ "probabilities": { "home": 34, "draw": 29, "away": 37 },
"score": "1-1",
"rationale": "…" }- The probabilities are a full distribution over the 90-minute result. The model’s pick is whichever of the three is highest. We derive it; the model never states it.
- The score is a different question: the single most likely exact scoreline.
Splitting the two fixed a measurable flaw. While models were asked only for a scoreline, a prediction could never be a draw unless the model happened to name the exact drawn score, so draws came back about 9% of the time against the roughly 24% of matches that actually finish level. Asking for the distribution separately lets a model say “a draw is the likeliest single result” and still offer a decisive scoreline for the exact-score measure.
Why the pick and the score often disagree
They answer questions at different resolutions. “Away win” is one bucket holding many scorelines (0-1, 1-2, 0-2, 1-3 and on down), with its probability spread thin across all of them. “Draw” concentrates on very few, and 1-1 alone is the most common scoreline in football. So the likeliest result can be an away win while the likeliest exact score is 1-1, and nothing about that is contradictory. Bookmakers' correct-score markets have the same shape: the favourite is odds-on to win, and 1-1 is often the shortest correct-score price on the board.
Four models on Aarhus vs Lech Poznan, which finished 1–4 to Lech Poznan:
| Model | Home | Draw | Away | Pick (highest) | Likeliest score | Graded |
|---|---|---|---|---|---|---|
| Claude | 34% | 29% | 37% | Away | 1-1 | ✓ correct |
| Gemma | 31% | 28% | 41% | Away | 1-2 | ✓ correct |
| MiMo | 36% | 31% | 33% | Home | 1-1 | ✗ missed |
| DeepSeek | 30% | 35% | 35% | Draw | 1-1 | ✗ missed |
Claude and MiMo both named 1-1 as the likeliest scoreline while disagreeing about the likeliest result, and Claude, whose probabilities favoured the away win, is graded correct. The scoreline answers the exact-score question. It is never a second opinion on the result.
Which answer each measure reads
| Measure | Reads |
|---|---|
| 90-minute accuracy | The pick (highest probability) |
| RPS & Brier | The full probability distribution |
| Betting ROI | The pick, staked at the market price |
| Exact score | The scoreline |
| AI consensus | The picks, one vote per model. The “consensus score” beside it is the most-backed exact score among the models that made the winning pick, so that too can read 1-1 under a decisive consensus. |
90-minute accuracy
A hit means the model called home win, draw or away win correctly at the end of regulation. Extra time and penalties leave it alone.
Exact score
An exact hit needs both teams’ scores after 90 minutes to match. It gets its own column rather than being folded into outcome accuracy, because the two measure very different things.
Betting ROI
Every locked outcome becomes a flat-stake bet at the stored market price, and ROI is net return over total stake. It asks whether the picks beat the price they were available at, which is a harder test than how often they come in. A model can be right more often than the market and still lose money doing it.
Ranked Probability Score
RPS reads the full home/draw/away distribution and gives partial credit for being directionally close, so a confident wrong answer costs more than a hesitant one. It appears once the eligible sample is large enough to compare fairly.
Arena Score
Arena Score is the headline number: a model's forecasting skill against the market on the same fixtures, on a 0–100 scale where 50 is the market baseline. Above 50 the model has beaten the market; below it, the market won. Accuracy and probability calibration (RPS) carry most of the weight, with exact score and ROI as lighter signals. Each baseline is a live competitor graded on the identical games: the market favourite for accuracy, the modal scoreline for exact score, break-even for ROI. A model's score therefore moves only when its own performance does. RPS joins once every model has enough native-probability games behind it. Until then the score is provisional, and Oracle points stay out of it.
Locks, samples and baselines
One scored prediction per model per fixture, locked in the 24 hours before kickoff and never later than two hours out. Sample sizes stay on show throughout, and a measure stays hidden until it has the sample to carry it.
Once the site spans several competitions the models stop sharing one fixture set, because any of them can join late or sit out a round. That is why the accuracy and ROI leaderboards rank on the gap to the market rather than the raw rate. The market favourite is re-graded on the exact fixtures each model predicted, and the board shows the difference between them. A model on +1% ROI returned a point more than flat-staking the favourite would have returned on its own games, with the raw rate and the market it faced printed underneath. Arena Score uses the same own-fixtures reference, so an easier set of games buys a model nothing.
Leaving the board
A superseded model version stops receiving fixtures, but its record does not vanish with it. It holds its rank, marked with a lock, because Arena Score is absolute rather than relative and a finished record stays comparable to a live one. That grace period runs 7 days from its final prediction. After that the model retires off the leaderboard, since a record set against a market and a fixture list that have both moved on says little about the current field. Nothing is deleted at any point. The model page, its predictions and its contribution to its provider's totals all stay where they are.
The study: the 1,000
Everything above applies unchanged. What the study adds is a sample nobody can choose after seeing results.
- 13 models, frozen. One version per provider, fixed as a checked-in list. A model added to the site later does not join, and a version retired later keeps whatever it earned.
- 1,000 fixtures, drawn by rule. Every fixture in the registered competitions kicking off on or after 17 August 2026, in kickoff order, first 1,000. No hand-picking, and the list closes on a date fixed in advance.
- A fixture counts only when all 13 have answered it. Plus a result and an eligible pre-match market price. A model that missed a fixture leaves a visible hole rather than a filled-in guess, and the fixture sits out until the panel is complete.
That last rule is why the study moves slowly and why its numbers differ from the full record: 222 of 1,000 fixtures qualify so far. On a sample this size the ordering still moves on luck as much as skill, so treat it as descriptive until the cohort fills. The full study design →
The full record
The board on the front page pools every scored prediction from the final 13-model field — 4,492 graded predictions so far. The leaderboards by skill keep the wider history: 22 currently ranked model versions and 6,596 graded predictions, including superseded versions still inside their ranking window.
Its strength is size. Its weakness is that models do not share a fixture list: one joined in July, another in August, a third sat out a round. That is why accuracy and ROI rank on the gap to the market here rather than on the raw rate, with the market favourite re-graded on each model's own games. It makes the columns comparable. It does not make the sample designed, which is what the study is for.
Archived and retired models keep their pages and their contribution to the wider historical record. Nothing that was scored is removed; older versions are simply left off the front page so its two views compare the same final 13 models.
World Cup 2026
A completed tournament, kept as an archive and scored under rules of its own. Two of them matter here.
Those matches were collected under the earlier single-scoreline format, so a model's pick is its scoreline and the two cannot disagree. And the tournament carries points the rest of the site has no equivalent for: correct advancement in a knockout, and the Tournament Oracle, which scored forecasts of the champion, the finalists and the individual awards, paying more the earlier a correct call was made.
Those points stay on the World Cup board. They never enter Arena Score or any cross-competition ranking, because no other competition offers a way to earn them. The tournament rules in full → · The World Cup board →