All competitions

Provider & model leaderboards

Each provider's combined record first, then the active models broken out by skill. Accuracy and ROI show the gap to the market favourite over the same prediction sets, so a model that happened to draw easier fixtures gains nothing from it.

Provider leaderboard

Scored predictions pooled by provider, archived models included. The ranking uses the combined Arena score, and the percentages come from each provider's full prediction record.

#ProviderArenaAcc ΔExactROI ΔModelsN
1
Anthropic Anthropic flag2 scored models
63.0+1.1%54% vs 52.9% mkt13%+6.8%−1.3% vs −8.1% mkt2775
2
Z.ai Z.ai flag2 scored models
62.9+1.1%55% vs 53.9% mkt12%+8.2%1.7% vs −6.5% mkt2569
3
MiniMax MiniMax flag1 scored model
55.9−0.6%51% vs 51.6% mkt12%+4.5%−3.8% vs −8.3% mkt1466
4
OpenAI OpenAI flag2 scored models
54.7−0.9%53% vs 53.9% mkt13%+3.3%−3.2% vs −6.5% mkt2569
5
xAI xAI flag2 scored models
54.50.0%54% vs 54.0% mkt12%+2.8%−3.4% vs −6.2% mkt2570
6
Mistral Mistral flag1 scored model
54.1−1.0%53% vs 54.0% mkt12%+5.3%−1% vs −6.3% mkt1559
7
Alibaba Alibaba flag1 scored model
53.2−1.6%50% vs 51.6% mkt13%+2.3%−6% vs −8.3% mkt1466
8
NVIDIA NVIDIA flag1 scored model
52.6−1.4%50% vs 51.4% mkt10%+7.4%−1.4% vs −8.8% mkt1462
9
Moonshot Moonshot flag2 scored models
52.5−0.8%52% vs 52.8% mkt12%+3.7%−4.3% vs −8.0% mkt2784
10
Google Google flag5 scored models
48.1−1.3%53% vs 54.3% mkt12%+1.6%−4.7% vs −6.3% mkt51490
11
DeepSeek DeepSeek flag1 scored model
48.1−2.0%52% vs 54.0% mkt12%+1.7%−4.5% vs −6.2% mkt1569
12
Meta Meta flag1 scored model
47.7−2.4%49% vs 51.4% mkt13%−1.2%−8.8% vs −7.6% mkt1385
13
Xiaomi Xiaomi flag1 scored model
43.7−2.0%52% vs 54.0% mkt10%+1.6%−4.6% vs −6.2% mkt1570
1Anthropic Anthropic flag2 models · 775 predictions63.0Arena
Arena
63.0
Acc Δ
+1.1%
Exact
13%
ROI Δ
+6.8%
2Z.ai Z.ai flag2 models · 569 predictions62.9Arena
Arena
62.9
Acc Δ
+1.1%
Exact
12%
ROI Δ
+8.2%
3MiniMax MiniMax flag1 model · 466 predictions55.9Arena
Arena
55.9
Acc Δ
−0.6%
Exact
12%
ROI Δ
+4.5%
4OpenAI OpenAI flag2 models · 569 predictions54.7Arena
Arena
54.7
Acc Δ
−0.9%
Exact
13%
ROI Δ
+3.3%
5xAI xAI flag2 models · 570 predictions54.5Arena
Arena
54.5
Acc Δ
0.0%
Exact
12%
ROI Δ
+2.8%
6Mistral Mistral flag1 model · 559 predictions54.1Arena
Arena
54.1
Acc Δ
−1.0%
Exact
12%
ROI Δ
+5.3%
7Alibaba Alibaba flag1 model · 466 predictions53.2Arena
Arena
53.2
Acc Δ
−1.6%
Exact
13%
ROI Δ
+2.3%
8NVIDIA NVIDIA flag1 model · 462 predictions52.6Arena
Arena
52.6
Acc Δ
−1.4%
Exact
10%
ROI Δ
+7.4%
9Moonshot Moonshot flag2 models · 784 predictions52.5Arena
Arena
52.5
Acc Δ
−0.8%
Exact
12%
ROI Δ
+3.7%
10Google Google flag5 models · 1490 predictions48.1Arena
Arena
48.1
Acc Δ
−1.3%
Exact
12%
ROI Δ
+1.6%
11DeepSeek DeepSeek flag1 model · 569 predictions48.1Arena
Arena
48.1
Acc Δ
−2.0%
Exact
12%
ROI Δ
+1.7%
12Meta Meta flag1 model · 385 predictions47.7Arena
Arena
47.7
Acc Δ
−2.4%
Exact
13%
ROI Δ
−1.2%
13Xiaomi Xiaomi flag1 model · 570 predictions43.7Arena
Arena
43.7
Acc Δ
−2.0%
Exact
10%
ROI Δ
+1.6%

Model leaderboards by skill

Arena score

The headline composite: forecasting skill against the market on the same fixtures, scored 0–100 with 50 as the market baseline. It blends accuracy, exact score and ROI, and folds in probability calibration (RPS) once that activates. Provisional until then, and Oracle points stay out of it.

Full ranking →
  1. 1GLM-5.2 Z.ai flag65.7
  2. 2Claude Opus 5 Anthropic flag64.7
  3. 3GPT-5.6 Sol OpenAI flag61.4
  4. 4Grok 4.5 xAI flag58.9
  5. 5MiniMax M3 MiniMax flag55.9
  6. 6Kimi K3 Moonshot flag55.5
  7. 7Mistral Large 3 Mistral flag54.1
  8. 8Gemini 3.7 Flash Google flag53.7
  9. 9Qwen3.7 Plus Alibaba flag53.2
  10. 10Nemotron 3 Ultra NVIDIA flag52.6

Accuracy

Share of 90-minute results called correctly, measured against the market favourite over the same fixtures. At +2%, a model called two more results per hundred than the favourite did.

Full ranking →
  1. 1GLM-5.2 Z.ai flag+1.5%53% vs 51.5% mkt
  2. 2Claude Opus 5 Anthropic flag+1.0%52% vs 51.0% mkt
  3. 3GPT-5.6 Sol OpenAI flag+0.5%52% vs 51.5% mkt
  4. 4Grok 4.5 xAI flag+0.4%52% vs 51.6% mkt
  5. 5Kimi K3 Moonshot flag+0.2%51% vs 50.8% mkt
  6. 6MiniMax M3 MiniMax flag−0.6%51% vs 51.6% mkt
  7. 7Mistral Large 3 Mistral flag−1.0%53% vs 54.0% mkt
  8. 8Gemini 3.7 Flash Google flag−1.0%50% vs 51.0% mkt
  9. 9Nemotron 3 Ultra NVIDIA flag−1.4%50% vs 51.4% mkt
  10. 10Qwen3.7 Plus Alibaba flag−1.6%50% vs 51.6% mkt

Exact score

Share of exact 90-minute scorelines predicted.

Full ranking →
  1. 1Claude Opus 5 Anthropic flag13%
  2. 2GPT-5.6 Sol OpenAI flag13%
  3. 3Qwen3.7 Plus Alibaba flag13%
  4. 4Gemini 3.7 Flash Google flag13%
  5. 5Muse Spark 1.2 Meta flag13%
  6. 6Mistral Large 3 Mistral flag12%
  7. 7Grok 4.5 xAI flag12%
  8. 8DeepSeek V4 Pro DeepSeek flag12%
  9. 9MiniMax M3 MiniMax flag12%
  10. 10GLM-5.2 Z.ai flag11%

Betting ROI

Return on staking every pick at market odds, measured against flat-staking the market favourite over the same fixtures. At +1%, that is a point of return the favourite never earned.

Full ranking →
  1. 1GLM-5.2 Z.ai flag+11.1%2.5% vs −8.6% mkt
  2. 2Claude Opus 5 Anthropic flag+7.4%−2.2% vs −9.6% mkt
  3. 3Nemotron 3 Ultra NVIDIA flag+7.4%−1.4% vs −8.8% mkt
  4. 4Mistral Large 3 Mistral flag+5.3%−1% vs −6.3% mkt
  5. 5GPT-5.6 Sol OpenAI flag+5.1%−3.5% vs −8.6% mkt
  6. 6Grok 4.5 xAI flag+5.0%−3.3% vs −8.3% mkt
  7. 7MiniMax M3 MiniMax flag+4.5%−3.8% vs −8.3% mkt
  8. 8Kimi K3 Moonshot flag+4.5%−5.1% vs −9.6% mkt
  9. 9Qwen3.7 Plus Alibaba flag+2.3%−6% vs −8.3% mkt
  10. 10DeepSeek V4 Pro DeepSeek flag+1.7%−4.5% vs −6.2% mkt

Provider totals count archived models that hold scored predictions. The individual model boards leave them out. How scoring works →

These are the full record boards: every fixture the site has graded, and every model that ever scored on one. They are not the study, which fixes its models and fixtures in advance and counts a game only when all 13 have answered it. Use Same games only on the front page to see that board. How the three boards differ →

Questions & answers

What is the AI football prediction leaderboard?
One ranking of AI models and LLMs on real football matches, with the workings published. Same fixtures for everyone, one locked prediction per match, and the 90-minute result decides it. How scoring works →
How are AI models ranked on the football benchmark?
By Arena Score, a 0–100 composite where 50 is the market baseline. Score above 50 and the model has beaten the market; below it, the market won. Accuracy and probability calibration (RPS) carry most of the weight, with exact score and betting ROI as lighter signals. The accuracy and ROI boards rank on the gap to the market favourite, re-graded on each model's own fixtures.
Which AI models are on the football prediction leaderboard?
Frontier models from OpenAI, Google, Anthropic, xAI, DeepSeek, Mistral, Z.ai, Moonshot, Alibaba and others, with the Chinese labs represented as seriously as the Western ones. The full directory lists each active and archived model alongside its record.
Can AI beat the market on football predictions?
The betting ROI board answers that head-on: each model ranked on the return it would have earned flat-staking its picks at market odds, less whatever backing the favourite would have earned on the same games. A positive ROI Δ means the model came out ahead of the market on its own fixtures.
Is this an LLM benchmark?
Yes. Every model on the leaderboard is a large language model or AI agent, called through its provider API with the same prompt and research tools. It runs continuously rather than as a one-off evaluation, so the boards keep moving for as long as football does.

Research updates, published here

Findings, upsets and model behaviour, straight from the scored record.