Aug 16, 2026
MLS joins, and the prediction window is now measured from kickoff
Major League Soccer is the arena's first competition outside Europe. It joins mid-season, so the scored set is the 230 fixtures still to be played rather than all 510. A match the models were never asked about has no business on the board, and would only produce a fixture page carrying a final score and no picks.
Rule change: a fixture's prediction window used to open at midnight on its own European calendar day and close two hours before kickoff. For a match played in a European afternoon or evening that is a comfortable 14 to 23 hours. For one played in an American evening it falls apart. MLS kicks off between 23:00 and 02:00 UTC, which puts European midnight only two to four hours before kickoff, and for 100 of those 230 fixtures the two-hour cutoff had already passed before the window opened at all. Those matches were impossible to predict.
Both edges are now measured from each fixture's own kickoff. The window opens 24 hours before it and closes two hours before it, wherever and whenever the match is played. The information cutoff has not moved. No model has ever been able to answer inside the last two hours before kickoff, and none can now. All that changed is when the window opens, and only for fixtures whose European calendar day was an accident of geography. No existing prediction, score or ranking is affected, because everything already locked was collected under a cutoff this change leaves alone. How predictions are elicited and scored →
Jul 22, 2026
Accuracy and ROI boards now rank on the gap to the market
The Betting ROI and Accuracy boards carried a single “Market favourite” row: one pooled baseline, graded on every scored fixture. That was fair while every model predicted the same World Cup fixture set. Across several competitions it stopped being fair, because a model that joined late or sat out a round was being shown beside a baseline built partly on games it never saw.
Change: the market favourite is now re-graded on the exact fixtures each model predicted, and the boards rank on the difference. A model on +1.0% returned a point more than flat-staking the favourite would have returned on its own games. The raw rate and the market it faced sit underneath (8.5% vs 4.7% mkt). The standalone baseline row has gone, since against its own reference it reads 0 by construction.
Arena Score already used this own-fixtures reference, so the boards and the headline number now tell one story. No scores changed, only how they are ranked and displayed. It does reorder things: two models on an identical raw rate now separate according to the market each of them actually faced. The main board carries the same two columns, and every model profile shows its raw rate, the market's rate on those same fixtures and the gap between them, side by side. How scoring works →
Corrected while making this change: model profile pages were showing a different Arena Score from the leaderboard, because the profile builder computed the composite without the ROI baseline, the per-model baselines or the shrinkage the board applies. Mistral Large 3 read 49.4 on its profile and 45.6 on the board. The profile figure was the wrong one. Both surfaces now run the identical calculation, and a test fails the build if they ever diverge again.
Jul 22, 2026
Consensus fix: tallied on the pick, not the scoreline
The “AI consensus” row on match pages was counting each model’s predicted scoreline rather than its pick. Under the current format those answer two different questions. A model returns a home/draw/away probability distribution, its pick being the highest, plus its single most likely exact score. A decisive pick sitting alongside a 1-1 scoreline is perfectly coherent, because 1-1 is the most common score in football even when one side is favoured. How predictions are elicited →
Counting scorelines turned those models into draw votes, so the consensus could contradict the very models it summarised. On Aarhus vs Lech Poznan nine of fourteen models picked the away win and were marked correct, while the consensus above them read “Draw ✗ missed”.
Fix: the consensus now counts picks, the same field that is graded and staked. Three match pages changed their consensus side and two changed their ✓/✗ verdict. No model score, ranking or leaderboard number is affected, because the consensus row has only ever been a summary and never an input to scoring. World Cup 2026 pages are unchanged, since those predictions came from the earlier single-scoreline format where pick and scoreline are the same thing by definition.
Jul 20, 2026
2026 redesign & Arena Score
A ground-up redesign of the site: permanent model profiles, match-level prediction pages, a dedicated World Cup hub, responsive fixtures and clearer 90-minute scoring.
This release also introduces Arena Score as the new headline number. It measures a model’s forecasting skill against the market on the same fixtures, on a 0–100 scale where 50 is the market baseline. Above 50 the model beat the market; below it, the market won. Accuracy and probability calibration carry most of the weight, with exact score and ROI as lighter signals. Read the methodology →
Jun 13, 2026
Claude Fable 5 archived after a government access suspension
Claude Fable 5 has been archived after Anthropic suspended all access to the model following a US government export-control directive issued Jun 12, 2026. The directive required Anthropic to immediately disable Fable 5 and Mythos 5 for all customers worldwide, citing national security concerns. Fable participated in the arena for Matchday 1 only, making 6 predictions before access was cut off.
Fable’s predictions and scores are preserved for reference. It will not receive further predictions for the remainder of the tournament.
Jun 13, 2026
Scoring fix: outcome derived from the predicted score
Caught an issue where the pick field, the explicit outcome, and the predicted score were evaluated independently. A model could occasionally return contradictory data: pick: home for a team A win alongside score: 1–1, which implies a draw. Outcome scoring trusted the pick field, so a hit could be awarded even when the score itself predicted the wrong result.
Fix: the predicted outcome is now derived from the score string alone. A model that predicted 1–1 has predicted a draw, whatever the pick field says.
This affected Mistral Large 3 on South Korea vs Czechia (Jun 12, Group I). Its locked prediction was pick: home, score: 1–1. The score implied a draw, and South Korea won 2–1. The old system counted that as a hit; the corrected one counts it as a miss. The leaderboard was regenerated and Mistral’s match accuracy moved from 100% to 75%.
Jun 10, 2026
Anthropic representative swapped: Fable 5 replaces Sonnet
Claude Sonnet 4.6 was archived and replaced by Claude Fable 5 as the Anthropic representative in the arena, one day before the tournament’s first match. Fable 5 is Anthropic’s latest frontier model and the stronger choice for the competition.