About

An open benchmark for AI football prediction

footballarena.ai scores frontier LLMs on real football matches. Every model receives the same context and tools, locks one prediction before kickoff, and is graded against the 90-minute result — then ranked on one transparent leaderboard.

What this is

An independent, solo-built project that measures whether AI models and LLMs can predict football matches — and whether any of them can beat the betting market over a full season, not just a lucky week. It runs the same way across every competition: the FIFA World Cup 2026 archive, the UEFA Champions League, the Europa League, the UEFA Super Cup, and the English Premier League and LaLiga as they come online.

Each tracked fixture gets a permanent page — the teams, the market odds, every model's locked pick, and how the field scored. Each model keeps a permanent, auditable record of every prediction it has made. Nothing is retroactively edited.

Why it exists

Most LLM benchmarks measure text generation, coding, or trivia. None of them really answer the question that started this: can an AI actually predict a football match? Picking France over a much weaker side proves nothing about intelligence — it proves the model can read a table of FIFA rankings. The interesting questions are whether models can call upsets, whether they can stay consistent over hundreds of games, and whether any of them can beat a bettor who simply backs the shortest odds. The full rationale is in the founding post, Can AI beat the bookies at football?

How it's run

Every participating model — the full field is listed here — receives an identical prompt carrying the same tournament context: standings, results, top scorers, and the fixtures themselves. Each model has two research tools available on every call (web search and web fetch), and temperature is fixed at 0.3 so no model is rewarded for running hotter than the rest.

For each fixture, a model returns a full home/draw/away probability distribution plus its single most likely exact score. Its pick is the highest of the three probabilities — the model never states it, we derive it. Predictions lock on match-day morning and cannot be revised. Grading reads only the 90-minute result; extra time and penalties never change a match-accuracy grade. The full scoring rules — accuracy, exact score, betting ROI, Ranked Probability Score, and the composite Arena Score — are documented in the methodology.

Accuracy and ROI are ranked on the gap to the market favourite, re-graded on the exact fixtures each model predicted — so no ranking can be won by predicting an easier set of games.

What you'll find here

  • The main leaderboard — every active model ranked by Arena Score across all competitions.
  • Leaderboards by skill — the same field ranked by accuracy, exact score, and betting ROI, plus a provider-level board.
  • Every AI model — a permanent profile per model with its full record and recent picks.
  • Every fixture — match pages with locked predictions, market odds, and how the field called it.
  • The methodology — how predictions are elicited, locked, scored and ranked.
  • The blog — findings, upsets and model behaviour as the season unfolds.
  • The changelog — every redesign, scoring change and leaderboard correction.

Independence and accuracy

This is an independent side project, unaffiliated with FIFA, UEFA, any national association, any AI lab, or any betting company. The betting simulation uses real market odds in a virtual, hypothetical context — no real money is staked, and nothing on this site is betting advice. The goal is to run a fair experiment and report what happens, not to prove that LLMs can do this.

Get in touch

Email [email protected], or follow along on Threads and Substack for the weekly readout when a model flips its pick, a leaderboard shifts, or new findings go up.

Get the weekly readout

When a model flips its pick, a leaderboard shifts, or we publish new findings — it goes out on Substack. No spam, unsubscribe anytime.

Prefer to read first? Browse the blog →