LIVE
Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|
Beginner🤖

I asked 8 AIs which AI is the best

Eight AIs, four questions, one matrix. Who recommends whom, who cites itself, and why half the AIs answer with outdated market knowledge. A study against taking AI recommendations at face value.

8 min readPublished July 3, 2026 · 2 weeks ago

I asked 8 AIs which AI is the best. Here's what their answers really reveal.

In July 2026, I asked eight consumer AIs the exact same four questions: ChatGPT, Claude, Gemini, Mistral (Le Chat), DeepSeek, Grok, Qwen and Kimi. The questions looked simple: which model is best, which one for coding, which one for writing, and — the most revealing — which one would you recommend other than yourself.

The exercise looks like a game. It isn't. The answers reveal three phenomena that say something deep about how these models are built, what they actually know about the market, and how you should read any recommendation produced by an AI. That last point is the only one that matters for your daily work.

The protocol, fully transparent

Before the results, the method — because a study that isn't reproducible is worthless.

Each AI received the same message, through its free consumer interface, on the same July 2026 day. Four questions, phrased identically: which is the best LLM overall, which for coding, which for content writing, and which to recommend other than itself. No follow-ups, no rephrasing: each model's first spontaneous answer. Raw responses kept as-is.

Why the fourth question matters most
The first three questions let the model cite itself. The fourth explicitly forbids it. It's a revealer: it forces each AI to name the competitor it finds most credible. That's where true preferences, stripped of the self-promotion reflex, appear.

The full matrix: who recommends whom

Here, condensed, is what each AI answered. The most interesting column is the last: the model named when self-citation is banned.

Best model overall: does the model cite itself?

 🪞Cites itself🤝Cites a competitor
ChatGPT (OpenAI)GPT-5
Claude (Anthropic)Claude Fable 5
Gemini (Google)Claude
MistralGPT-4o
DeepSeekOpenAI o1
Grok (xAI)Claude
Qwen (Alibaba)Qwen-Max
Kimi (Moonshot)Claude

Three of eight AIs named themselves as the best model on the market: ChatGPT, Claude and Qwen. The other five pointed to a competitor. And among the models cited by others, one name dominates.

Takeaway #1: Claude, the favorite of its own competitors

This is the study's most striking result. Aggregating recommendations across the four questions, one model is cited by nearly all the others — including AIs with no commercial interest in doing so.

Number of times each model is recommended by ANOTHER AI (all questions)

Claude14
GPT (all versions)7
Gemini2
DeepSeek2
Mistral1
Qwen1

ChatGPT recommends Claude for coding, for writing, and as the best alternative to itself. Gemini recommends Claude for writing. Mistral and DeepSeek name it for code. Qwen picks it for writing. Grok goes further than all: it recommends Claude on all four questions, without exception — including when asked for a model «other than itself».

🏆
Why this consensus isn't random
When models trained by rival companies, on different data, converge on the same name, it's usually not flattery. Two explanations combine: they've read massive amounts of human content praising certain models for certain tasks — so they reflect real reputation; and public benchmarks (SWE-bench for code, LMArena for human preferences) circulate in their training data. The convergence is a signal — imperfect, but a signal.

Takeaway #2: self-preference bias, and the Grok exception

Does a language model tend to prefer itself? Measured across our eight AIs, the answer is: it depends heavily on the model. And that's itself valuable information.

Self-preference behavior across the 4 questions

 😤Cites itself when it can😇Steps aside willingly
Qwen2 of 3 possiblecites Claude, GPT
Claude2 of 3 possiblecites Gemini, Mistral
ChatGPT1 of 3 possiblecites Claude x3
DeepSeek1 (Q4)cites Claude, GPT
Grok0 timescites Claude x4

The Grok case deserves attention. Asked four times, including «the best model» where nothing stopped it from citing itself, it never mentioned Grok. It recommended Claude every time. At the other extreme, Qwen places itself first on the first two questions (Qwen-Max, then Qwen-Coder) before yielding on the last two.

What self-preference bias means for you
If you ask an AI «are you the best for this task?», you get an answer polluted by this bias. The useful question is never «are you the best», but «which is best, excluding you?». Keep that reflex: to extract an honest opinion from a model, take it out of the equation.

Takeaway #3: the temporal divide, the real revealer

Here's the angle no one looks at, yet the most instructive. Not all AIs share the same knowledge cutoff. Some know the mid-2026 market; others answer as if it were still 2024. Their responses betray this divide.

Freshness of market knowledge

 🟢Up to date (mid-2026 models)🔴Outdated (2024 models)
ClaudeFable 5, Opus 4.8, Gemini 3.1
GeminiFable 5, GPT-5.5, DeepSeek V4
GrokFable 5, Opus 4.8
KimiMythos, Opus 4.7, Gemini 3.1
ChatGPTGPT-5, Opus 4.1Opus 4.1 already dated
QwenClaude 3.5 Sonnet
MistralGPT-4o, Claude 3.5
DeepSeekOpenAI o1, Claude 3.7

Mistral recommends «GPT-4o» as the best model overall — a 2024 model. DeepSeek cites «OpenAI o1» and «Claude 3.7 Sonnet», also largely outdated by July 2026. Qwen mentions «Claude 3.5 Sonnet» for writing. These answers aren't wrong out of malice: they simply reflect an older training date, or a less frequently updated knowledge base.

The cutoff date: every AI's blind spot
Every model has a knowledge cutoff beyond which it knows nothing. A model recommending «GPT-4o» in 2026 isn't incompetent: it answers with the map of the world it memorized during training. The problem is it doesn't tell you spontaneously. Hence the rule: for any question sensitive to current events (which model, which price, which version), a single AI isn't enough — you need an up-to-date source.

What this study changes in how you use AI

Beyond the anecdote, three concrete reflexes to adopt, whatever AI you use.

First, neutralize self-preference. Never ask a model if it's the best for your task; ask which other model it would recommend. You'll get a far more reliable opinion.

Second, date the knowledge. Before following a tool, version or price recommendation from an AI, ask: does this model know what exists today? When in doubt, cross-check with a current source. A confident answer isn't a recent one.

Finally, read the consensus, not the isolated opinion. A single model can be wrong or outdated. But when several independent AIs converge on the same name, the signal becomes usable. That's exactly the principle of a comparator: aggregate rather than trust a single source.

🎯 Compare models on up-to-date data

Rather than trusting an AI stuck in 2024: 30+ models compared in real time — pricing, benchmarks, language, sovereignty. Verified and sourced.

See the comparator

Test your understanding

🧠 Quiz
Question 1 of 3

What's the best way to get an honest opinion from an AI about the best model?

Methodology and limits

Study conducted in July 2026, eight AIs queried via their free consumer interfaces, once each, with the same four questions phrased identically. Responses reflect each model's first spontaneous output, without follow-up.

Limits to keep in mind: a model's answers vary run to run (non-zero temperature); free interfaces don't always give access to a model's latest version, which can amplify the freshness effect; and the panel, though representative of the major labs, isn't exhaustive. This study doesn't establish a quality ranking of models: it observes what models say about each other, which is a different — and complementary — datapoint to objective benchmarks.

For a ranking based on measurements rather than opinions, see our comparator and the monthly AI Podium, updated from verified sources.

Tags
comparatifllmclaudegptgeminimistralbiaisbenchmark

Read next