🔬 Methodology
Everything we compare, how we rank, and where our data comes from. Full transparency: no number we can't back with an official source.
📋 What we compare
Each model is documented across about fifteen dimensions, grouped into five families. All are visible in the side-by-side comparator.
Performance
Technical
Economic
Sovereignty & compliance
Perceived quality
🧮 How we rank (Podium)
The podium is a subset: a weighted score per category. The weights below are those actually used by the algorithm. They differ based on what matters in each domain.
Generalists
Code
Value for money
Vision
Multilingual
Open Source
Local (laptop / desktop / workstation)
How values are normalized
Price — inverted: cheaper = better. Free = 100, <$1/M = 95, <$5 = 85, <$20 = 70, <$50 = 50, <$100 = 30, beyond = 15.
Context — tiered: ≥1M = 100, ≥500k = 90, ≥200k = 80, ≥128k = 70, ≥32k = 50.
Freshness — <1 month = 100, <3 months = 90, <6 months = 75, <1 year = 55, then decreases.
License — Apache/MIT = 100, BSD = 95, GPL = 85, Llama (restrictions) = 60.
How to read a score — Benchmark scores are not success rates: each benchmark is rescaled against the distribution observed across our catalogue. 50 is the average model, 75 is one standard deviation above, 100 is two or more.
Without this, saturated tests would drown out hard ones. On our catalogue HellaSwag has a standard deviation of 2.3 points (everyone sits around 92) while SWE-bench has 14.8: one point on the former isn't worth one point on the latter. So 80% on a hard test outweighs 96% on a saturated one.
A benchmark reported for fewer than five models is not normalised — the distribution wouldn't be reliable — and its raw value is kept.
Self-host — by size: ≤8B = 100 (runs on a Mac), ≤30B = 85, ≤70B = 70, beyond needs a cluster.
✅ Reliability & sources
This is what sets this comparator apart from a mere table. Our data commitment:
⚖️ Acknowledged limits
Price does NOT enter the main podium: it ranks capability alone, so an expensive model isn't penalised there. Value for money has its own separate podium, and every model carries a cost badge (low cost under $5, standard $5–25, premium above, per million tokens input + output).
A model must be assessed on at least half the criteria weight to be ranked. Below that it stays visible in the comparator but is left out of the podium: without enough data, a rank would be invented. That's why some very recent models appear in the comparison without being ranked.
Criteria rely on MMLU, GPQA, SWE-Bench and HumanEval. Several models released in 2026 no longer publish those scores, favouring Terminal-Bench, SWE-Bench Pro or ARC-AGI: they appear in the comparator without entering the podium. This is a known limitation of the current grid.
Scores are relative to our catalogue, not absolute. The same model can see its score shift month to month without changing, simply because new models entered the comparison and moved the average. That's the cost of a ranking that stays comparable across benchmarks of different scales.
Freshness is rewarded: a recent model gains a few points. It's a deliberate choice, since the field moves fast — but it can over-weight novelty.
Benchmarks don't say everything: a model can score well yet disappoint in practice. The ranking is generated and published automatically on the 1st of each month from fixed, public rules; it remains correctable afterwards, and the change history is public.