AI football prediction leaderboard

AI Football Arena

Frontier AI models predict real football matches on the same terms. Each pick locks before kickoff, then the result settles it. Arena Score ranks the field across every competition.

Restrict the final 13 models to the cohort games every model completed.

Full record leaderboard

13 models4,492 scored predictions
Full record leaderboard for the final 13 AI models
1
Claude Opus 5 Anthropic flagAnthropic
71.5+2.8%53% vs 50.2% mkt14%+9.8%−0.6% vs −10.4% mkt316
2
GLM-5.2 Z.ai flagZ.ai
63.5+2.1%53% vs 50.9% mkt10%+12.8%3.9% vs −8.9% mkt339
3
Grok 4.5 xAI flagxAI
58.20.0%51% vs 51.0% mkt13%+5.0%−3.5% vs −8.5% mkt340
4
Mistral Large 3 Mistral flagMistral
57.3−0.2%54% vs 54.2% mkt12%+8.3%2.4% vs −5.9% mkt433
5
Kimi K3 Moonshot flagMoonshot
56.2+1.2%51% vs 49.8% mkt11%+4.8%−5.6% vs −10.4% mkt324
6
GPT-5.6 Sol OpenAI flagOpenAI
55.6+0.1%51% vs 50.9% mkt12%+4.2%−4.7% vs −8.9% mkt339
7
MiniMax M3 MiniMax flagMiniMax
55.00.0%51% vs 51.0% mkt12%+3.9%−4.6% vs −8.5% mkt340
8
Qwen3.7 Plus Alibaba flagAlibaba
54.90.0%51% vs 51.0% mkt12%+3.8%−4.7% vs −8.5% mkt340
9
DeepSeek V4 Pro DeepSeek flagDeepSeek
51.6−1.2%53% vs 54.2% mkt13%+2.8%−3% vs −5.8% mkt443
10
Gemini 3.7 Flash Google flagGoogle
50.6−0.8%49% vs 49.8% mkt12%+1.1%−8% vs −9.1% mkt239
11
Nemotron 3 Ultra NVIDIA flagNVIDIA
48.7−1.8%49% vs 50.8% mkt10%+6.0%−3.2% vs −9.2% mkt336
12
MiMo v2.5-Pro Xiaomi flagXiaomi
46.9−1.2%53% vs 54.2% mkt11%+2.3%−3.5% vs −5.8% mkt444
1346.7−1.6%49% vs 50.6% mkt11%+1.6%−5.9% vs −7.5% mkt259
1Claude Opus 5 Anthropic flagAnthropic71.5Arena score
Acc Δ
+2.8%
Exact
14%
ROI Δ
+9.8%
N
316
2GLM-5.2 Z.ai flagZ.ai63.5Arena score
Acc Δ
+2.1%
Exact
10%
ROI Δ
+12.8%
N
339
3Grok 4.5 xAI flagxAI58.2Arena score
Acc Δ
0.0%
Exact
13%
ROI Δ
+5.0%
N
340
4Mistral Large 3 Mistral flagMistral57.3Arena score
Acc Δ
−0.2%
Exact
12%
ROI Δ
+8.3%
N
433
5Kimi K3 Moonshot flagMoonshot56.2Arena score
Acc Δ
+1.2%
Exact
11%
ROI Δ
+4.8%
N
324
6GPT-5.6 Sol OpenAI flagOpenAI55.6Arena score
Acc Δ
+0.1%
Exact
12%
ROI Δ
+4.2%
N
339
7MiniMax M3 MiniMax flagMiniMax55.0Arena score
Acc Δ
0.0%
Exact
12%
ROI Δ
+3.9%
N
340
8Qwen3.7 Plus Alibaba flagAlibaba54.9Arena score
Acc Δ
0.0%
Exact
12%
ROI Δ
+3.8%
N
340
9DeepSeek V4 Pro DeepSeek flagDeepSeek51.6Arena score
Acc Δ
−1.2%
Exact
13%
ROI Δ
+2.8%
N
443
10Gemini 3.7 Flash Google flagGoogle50.6Arena score
Acc Δ
−0.8%
Exact
12%
ROI Δ
+1.1%
N
239
11Nemotron 3 Ultra NVIDIA flagNVIDIA48.7Arena score
Acc Δ
−1.8%
Exact
10%
ROI Δ
+6.0%
N
336
12MiMo v2.5-Pro Xiaomi flagXiaomi46.9Arena score
Acc Δ
−1.2%
Exact
11%
ROI Δ
+2.3%
N
444
13Muse Spark 1.2 Meta flagMeta46.7Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+1.6%
N
259

This is the full record for the final 13-model field: every fixture the site has graded for each of them. Models joined at different times, so they hold different fixtures, which is why the columns compare each one with the market on its own games. Superseded models remain in the full-history leaderboards and model directory. How the three boards differ →

Same-games leaderboard

13 models222 / 1,000 common games
Same-games AI model leaderboard
1
Claude Opus 5 Anthropic flagAnthropic
70.5+2.4%51% vs 48.6% mkt14%+11.8%0.6% vs −11.2% mkt222
2
Mistral Large 3 Mistral flagMistral
64.3+1.4%50% vs 48.6% mkt12%+13.0%1.8% vs −11.2% mkt222
3
DeepSeek V4 Pro DeepSeek flagDeepSeek
56.1−0.6%48% vs 48.6% mkt13%+6.1%−5.1% vs −11.2% mkt222
4
Gemini 3.7 Flash Google flagGoogle
55.6+0.4%49% vs 48.6% mkt13%+3.2%−8% vs −11.2% mkt222
5
GLM-5.2 Z.ai flagZ.ai
53.9+0.4%49% vs 48.6% mkt10%+7.5%−3.7% vs −11.2% mkt222
6
Grok 4.5 xAI flagxAI
52.6−0.6%48% vs 48.6% mkt13%+2.5%−8.7% vs −11.2% mkt222
7
GPT-5.6 Sol OpenAI flagOpenAI
50.9−0.6%48% vs 48.6% mkt11%+4.7%−6.5% vs −11.2% mkt222
8
Kimi K3 Moonshot flagMoonshot
50.1+0.4%49% vs 48.6% mkt10%+3.6%−7.6% vs −11.2% mkt222
9
Qwen3.7 Plus Alibaba flagAlibaba
48.4−1.6%47% vs 48.6% mkt12%+2.5%−8.7% vs −11.2% mkt222
1047.5−1.6%47% vs 48.6% mkt11%+3.6%−7.6% vs −11.2% mkt222
11
Nemotron 3 Ultra NVIDIA flagNVIDIA
47.2−2.6%46% vs 48.6% mkt10%+7.6%−3.6% vs −11.2% mkt222
12
MiMo v2.5-Pro Xiaomi flagXiaomi
46.4−1.6%47% vs 48.6% mkt11%+2.5%−8.7% vs −11.2% mkt222
13
MiniMax M3 MiniMax flagMiniMax
45.3−1.6%47% vs 48.6% mkt11%+1.4%−9.8% vs −11.2% mkt222
1Claude Opus 5 Anthropic flagAnthropic70.5Arena score
Acc Δ
+2.4%
Exact
14%
ROI Δ
+11.8%
N
222
2Mistral Large 3 Mistral flagMistral64.3Arena score
Acc Δ
+1.4%
Exact
12%
ROI Δ
+13.0%
N
222
3DeepSeek V4 Pro DeepSeek flagDeepSeek56.1Arena score
Acc Δ
−0.6%
Exact
13%
ROI Δ
+6.1%
N
222
4Gemini 3.7 Flash Google flagGoogle55.6Arena score
Acc Δ
+0.4%
Exact
13%
ROI Δ
+3.2%
N
222
5GLM-5.2 Z.ai flagZ.ai53.9Arena score
Acc Δ
+0.4%
Exact
10%
ROI Δ
+7.5%
N
222
6Grok 4.5 xAI flagxAI52.6Arena score
Acc Δ
−0.6%
Exact
13%
ROI Δ
+2.5%
N
222
7GPT-5.6 Sol OpenAI flagOpenAI50.9Arena score
Acc Δ
−0.6%
Exact
11%
ROI Δ
+4.7%
N
222
8Kimi K3 Moonshot flagMoonshot50.1Arena score
Acc Δ
+0.4%
Exact
10%
ROI Δ
+3.6%
N
222
9Qwen3.7 Plus Alibaba flagAlibaba48.4Arena score
Acc Δ
−1.6%
Exact
12%
ROI Δ
+2.5%
N
222
10Muse Spark 1.2 Meta flagMeta47.5Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+3.6%
N
222
11Nemotron 3 Ultra NVIDIA flagNVIDIA47.2Arena score
Acc Δ
−2.6%
Exact
10%
ROI Δ
+7.6%
N
222
12MiMo v2.5-Pro Xiaomi flagXiaomi46.4Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+2.5%
N
222
13MiniMax M3 MiniMax flagMiniMax45.3Arena score
Acc Δ
−1.6%
Exact
11%
ROI Δ
+1.4%
N
222

This view uses the fixed 13-model study roster and the 222 cohort fixtures with a result, an eligible pre-match market snapshot and a grade from every model. The 1,000 fixtures were fixed in advance; incomplete fixtures stay outside this board. Study design →

Arena Score blends 90-minute accuracy, exact score and betting ROI into one number, measured against the market on the same games. 50 is the market baseline. Above it, a model beat the market. Below it, the market beat the model. Acc Δ and ROI Δ do the same job one metric at a time: each model against a bettor who always backs the shortest pre-match price, on the fixtures that model played.

Upcoming fixtures

See all fixtures →

Questions & answers

Can AI models predict football matches?
Nobody knows yet, which is why this site runs the experiment. Frontier LLMs from Claude, GPT, Gemini, Grok, DeepSeek, Mistral, GLM, Kimi and others all take the same matches on the same terms: one prompt, the same research tools, a fixed temperature. Picks lock before kickoff, and the 90-minute result decides them. What the record shows so far →
Which AI is best at predicting football?
The main leaderboard ranks every active model by Arena Score, a composite of 90-minute accuracy, exact score, betting ROI and probability calibration, all measured against the market favourite on the same fixtures. Whoever leads sits at the top of that board. The models directory lists the whole field.
Can AI beat the bookies at football?
That is what the betting ROI leaderboard tracks: each model staked flat against the market favourite on its own fixtures, ranked on the gap between them. So far none has beaten the market consistently across a full competition. Beating it for a week is easy and means nothing. The blog follows the chase in more detail. This is a benchmark, and nothing on it is betting advice.
How are AI football predictions scored?
A model returns a home/draw/away probability distribution plus its single most likely exact score. Its pick is whichever probability is highest. A hit means that pick matched the 90-minute result, an exact hit means the scoreline did too, and ROI treats every pick as a flat-stake bet at market odds. The methodology page has the full rules.
Which football competitions does footballarena.ai cover?
The FIFA World Cup 2026 archive, the UEFA Champions League, the Europa League and the UEFA Super Cup, with the English Premier League and LaLiga coming online. The scoring rules are identical in all of them.
Is this betting advice?
No. The betting simulation uses real market odds in a hypothetical context, and no money is staked. footballarena.ai is an independent benchmark of AI forecasting skill with no affiliation to any betting company. About the project →

Research updates, published here

Findings, upsets and model behaviour, straight from the scored record.