About
An open benchmark for AI football prediction
footballarena.ai scores frontier LLMs on real football matches. Same context, same tools, one prediction per match, locked before kickoff. The 90-minute result settles it, and a single leaderboard carries them all.
What this is
One person, one question. Can an AI predict a football match, and can it keep doing it for a season rather than for a lucky week? The betting market is the thing to beat, and the same rules apply wherever the site goes: the FIFA World Cup 2026 archive, the UEFA Champions League, the Europa League, the UEFA Super Cup, with the English Premier League and LaLiga joining as they come online.
Every tracked fixture keeps a page of its own: the teams, the market odds, each model's locked pick, and how the field scored. Every model keeps the full record of what it called. Nobody goes back and tidies it afterwards.
Why it exists
Most LLM benchmarks measure text generation, coding or trivia. None of them get near the question that started this one: can an AI actually predict a football match? Picking France over a much weaker side says nothing about intelligence. It says the model can read a FIFA ranking table. The harder questions come after that. Can it call an upset? Can it still be doing this well three hundred games later? And can it beat a bettor who does nothing cleverer than back the shortest odds every week? The founding post works through the reasoning: Can AI beat the bookies at football?
How it's run
Every model in the field (listed here) gets an identical prompt carrying the same competition context: standings, recent results, top scorers and the fixtures themselves. Two research tools come with every call, web search and web fetch. Temperature sits at 0.3 for everyone, so nobody profits from running hotter than the rest.
For each fixture a model returns a full home/draw/away probability distribution and its single most likely exact score. Its pick is whichever of the three probabilities runs highest. The model never names it; we read it off. Predictions lock in the 24 hours before kickoff, never later than two hours out, and nothing can be revised after that. Grading reads the 90-minute result and nothing else, so extra time and penalties leave a match-accuracy grade alone. The methodology covers the rest: accuracy, exact score, betting ROI, Ranked Probability Score, and how the composite Arena Score gets built.
Accuracy and ROI rank on the gap to the market favourite, re-graded on the exact fixtures each model played. An easier set of games therefore buys a model nothing.
Where to look
- The main leaderboard, where every active model is ranked by Arena Score across all competitions.
- Leaderboards by skill: the same field re-ranked by accuracy, by exact score and by betting ROI, with a provider-level board alongside.
- A profile for each model, carrying its full record and its recent picks.
- A page for each fixture, with the locked predictions, the market odds and how the field called it.
- The methodology, covering how a prediction is asked for, locked, scored and ranked.
- Research updates on findings, upsets and model behaviour drawn from the scored record.
- The changelog, listing redesigns, scoring changes and leaderboard corrections.
Independence
A side project with no affiliation to FIFA, UEFA, any national association, any AI lab or any betting company. The betting simulation uses real market odds in a hypothetical context. No money is staked, and nothing here is betting advice. The aim is to run the experiment fairly and publish whatever comes out of it, up to and including the finding that none of these models can do this.
Get in touch
Findings and updates are published here on footballarena.ai. For anything else, email [email protected]. New work is also linked from Threads, and the older Substack archive is still up.