Research study
Study design & methodology
Fixtures and panel were both fixed in advance, so nothing about the sample can be chosen after the fact. Every comparison then runs on the same completed games.
Hypothesis
Given identical structured match context and no tools beyond the declared ones, frontier LLMs differ from each other in football-prediction skill, on calibration and accuracy together, and those differences hold across a fixed cohort of 1,000 fixtures. The claim can be shown false: if the ranking on the common cohort turns out to be statistically indistinguishable from noise once all 1,000 fixtures are in, the hypothesis has failed.
Cohort definition
"The 1,000" is a single fixture list, checked into the repository and drawn by rule rather than picked by hand. Take every fixture in the registered competitions kicking off on or after 17 August 2026, order it by kickoff time and keep the first thousand, with competition and fixture ids breaking ties between simultaneous kickoffs. That start date is the first kickoff by which all 13 research models had a graded prediction, which keeps out any fixture a panel member could never have answered. The draw closes on 8 November 2026, when the thousandth fixture kicks off. Changes to the live fixture catalog do not touch this list.
A fixture reaches the research results once it has a final score, an eligible pre-match market snapshot and a graded prediction from all 13 models. A model run that never happened stays visible as a coverage gap; we do not fill it in with an estimate. Until the panel is complete on a fixture, that fixture moves no ranking here.
Model roster
The panel is frozen at one model per provider. These 13 versions are all of it:
- Anthropic → Claude Opus 5
- DeepSeek → DeepSeek V4 Pro
- Google → Gemini 3.7 Flash
- Z.ai → GLM-5.2
- OpenAI → GPT-5.6 Sol
- xAI → Grok 4.5
- Moonshot → Kimi K3
- Xiaomi → MiMo v2.5-Pro
- MiniMax → MiniMax M3
- Mistral → Mistral Large 3
- Meta → Muse Spark 1.2
- NVIDIA → Nemotron 3 Ultra
- Alibaba → Qwen3.7 Plus
Older versions, and models added after the protocol was frozen, stay in the live site record. Neither joins this panel or its boards.
Live coverage
The front page uses these same 13 models in both views. Its default full record keeps every graded prediction each one has made; Same games only narrows them to the common cohort, so a growing schedule or a missing prediction cannot leave one model measured on a different sample from another. Older model versions remain in the dedicated leaderboards and model directory.
Scoring mechanics
The grading rules for Arena Score, accuracy, exact score, betting ROI and calibration all live on the general methodology page. The same-games view applies them to its fixed common panel and nowhere else.