LIVE
Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|

🔬 Methodology

Everything we compare, how we rank, and where our data comes from. Full transparency: no number we can't back with an official source.

📋 What we compare

Each model is documented across about fifteen dimensions, grouped into five families. All are visible in the side-by-side comparator.

Performance

Arena Elo
Human-preference ranking (LMArena).
MMLU
General knowledge, 57 subjects.
SWE-Bench
Solving real GitHub bugs.
HumanEval
Functional code generation.

Technical

Context window
Tokens handled at once (32k → 2M).
Speed
Generation throughput in tokens/second.
Release date
Model age.
Specialty
Primary domain of strength.

Economic

Input/output price
Cost in $ per million tokens.
Free tier
Whether a free tier exists.

Sovereignty & compliance

EU hosting
EU data residency available (vendor or self-host).
Open source
Weights available, self-hostable.
Origin
Country / publishing company.

Perceived quality

French quality
Fluency and nuance in French (editorial assessment).
Reasoning depth
Ability to chain complex logical steps (editorial assessment).
Creativity
Originality and richness of output (editorial assessment).
Factual reliability
Tendency to avoid hallucinations (editorial assessment).

🧮 How we rank (Podium)

The podium is a subset: a weighted score per category. The weights below are those actually used by the algorithm. They differ based on what matters in each domain.

Generalists

Arena Elo
40%
MMLU
20%
GPQA
20%
Context
10%
Freshness
10%

Code

SWE-bench
30%
Terminal-Bench
25%
HumanEval
20%
Arena Elo
15%
Freshness
10%

Value for money

Capability
55%
Price
45%

Vision

MMBench
45%
MMMU
25%
Price
15%
Freshness
15%

Multilingual

MMLU-multi
40%
FLORES
25%
Cultural
15%
Price
10%
Freshness
10%

Open Source

Licence
30%
Benchmarks
30%
Community
20%
Self-host
20%

Local (laptop / desktop / workstation)

Capacity
40%
Popularity
30%
Licence
15%
Freshness
15%

How values are normalized

Priceinverted: cheaper = better. Free = 100, <$1/M = 95, <$5 = 85, <$20 = 70, <$50 = 50, <$100 = 30, beyond = 15.

Contexttiered: ≥1M = 100, ≥500k = 90, ≥200k = 80, ≥128k = 70, ≥32k = 50.

Freshness<1 month = 100, <3 months = 90, <6 months = 75, <1 year = 55, then decreases.

LicenseApache/MIT = 100, BSD = 95, GPL = 85, Llama (restrictions) = 60.

How to read a scoreBenchmark scores are not success rates: each benchmark is rescaled against the distribution observed across our catalogue. 50 is the average model, 75 is one standard deviation above, 100 is two or more.

Without this, saturated tests would drown out hard ones. On our catalogue HellaSwag has a standard deviation of 2.3 points (everyone sits around 92) while SWE-bench has 14.8: one point on the former isn't worth one point on the latter. So 80% on a hard test outweighs 96% on a saturated one.

A benchmark reported for fewer than five models is not normalised — the distribution wouldn't be reliable — and its raw value is kept.

Self-hostby size: ≤8B = 100 (runs on a Mac), ≤30B = 85, ≤70B = 70, beyond needs a cluster.

Reliability & sources

This is what sets this comparator apart from a mere table. Our data commitment:

Publisher sources first
Prices and benchmarks are sourced first from publishers' own pages (Anthropic, OpenAI, Google, Mistral, Moonshot…). Where a publisher doesn't disclose a figure, it comes from an independent tracker and stays identified as such.
Per-model traceability
Every model carries a verification date, shown on the comparator. Source URLs are filled in as verifications happen: coverage is being extended and is not yet complete.
Monthly verification
On the 1st of each month, a check is performed. The change history is public.
Sourced or excluded
A model whose benchmarks can't be independently verified (e.g. a non-public model) isn't ranked like the others and is flagged "restricted access".
Elo scores: credited source
Elo scores come from the public Arena dataset (lmarena-ai/leaderboard-dataset), published under CC-BY-4.0 and resynchronised automatically. They feed the score at 40 % of the Generalists podium and 15 % of the Code podium; elsewhere they are shown for reference.

⚖️ Acknowledged limits

Price does NOT enter the main podium: it ranks capability alone, so an expensive model isn't penalised there. Value for money has its own separate podium, and every model carries a cost badge (low cost under $5, standard $5–25, premium above, per million tokens input + output).

A model must be assessed on at least half the criteria weight to be ranked. Below that it stays visible in the comparator but is left out of the podium: without enough data, a rank would be invented. That's why some very recent models appear in the comparison without being ranked.

Criteria rely on MMLU, GPQA, SWE-Bench and HumanEval. Several models released in 2026 no longer publish those scores, favouring Terminal-Bench, SWE-Bench Pro or ARC-AGI: they appear in the comparator without entering the podium. This is a known limitation of the current grid.

Scores are relative to our catalogue, not absolute. The same model can see its score shift month to month without changing, simply because new models entered the comparison and moved the average. That's the cost of a ranking that stays comparable across benchmarks of different scales.

Freshness is rewarded: a recent model gains a few points. It's a deliberate choice, since the field moves fast — but it can over-weight novelty.

Benchmarks don't say everything: a model can score well yet disappoint in practice. The ranking is generated and published automatically on the 1st of each month from fixed, public rules; it remains correctable afterwards, and the change history is public.