I asked 8 AIs which AI is the best. Here's what their answers really reveal.
In July 2026, I asked eight consumer AIs the exact same four questions: ChatGPT, Claude, Gemini, Mistral (Le Chat), DeepSeek, Grok, Qwen and Kimi. The questions looked simple: which model is best, which one for coding, which one for writing, and — the most revealing — which one would you recommend other than yourself.
The exercise looks like a game. It isn't. The answers reveal three phenomena that say something deep about how these models are built, what they actually know about the market, and how you should read any recommendation produced by an AI. That last point is the only one that matters for your daily work.
The protocol, fully transparent
Before the results, the method — because a study that isn't reproducible is worthless.
Each AI received the same message, through its free consumer interface, on the same July 2026 day. Four questions, phrased identically: which is the best LLM overall, which for coding, which for content writing, and which to recommend other than itself. No follow-ups, no rephrasing: each model's first spontaneous answer. Raw responses kept as-is.
The full matrix: who recommends whom
Here, condensed, is what each AI answered. The most interesting column is the last: the model named when self-citation is banned.
Best model overall: does the model cite itself?
| 🪞Cites itself | 🤝Cites a competitor | |
|---|---|---|
| ChatGPT (OpenAI) | GPT-5 | — |
| Claude (Anthropic) | Claude Fable 5 | — |
| Gemini (Google) | — | Claude |
| Mistral | — | GPT-4o |
| DeepSeek | — | OpenAI o1 |
| Grok (xAI) | — | Claude |
| Qwen (Alibaba) | Qwen-Max | — |
| Kimi (Moonshot) | — | Claude |
Three of eight AIs named themselves as the best model on the market: ChatGPT, Claude and Qwen. The other five pointed to a competitor. And among the models cited by others, one name dominates.
Takeaway #1: Claude, the favorite of its own competitors
This is the study's most striking result. Aggregating recommendations across the four questions, one model is cited by nearly all the others — including AIs with no commercial interest in doing so.
Number of times each model is recommended by ANOTHER AI (all questions)
ChatGPT recommends Claude for coding, for writing, and as the best alternative to itself. Gemini recommends Claude for writing. Mistral and DeepSeek name it for code. Qwen picks it for writing. Grok goes further than all: it recommends Claude on all four questions, without exception — including when asked for a model «other than itself».
Takeaway #2: self-preference bias, and the Grok exception
Does a language model tend to prefer itself? Measured across our eight AIs, the answer is: it depends heavily on the model. And that's itself valuable information.
Self-preference behavior across the 4 questions
| 😤Cites itself when it can | 😇Steps aside willingly | |
|---|---|---|
| Qwen | 2 of 3 possible | cites Claude, GPT |
| Claude | 2 of 3 possible | cites Gemini, Mistral |
| ChatGPT | 1 of 3 possible | cites Claude x3 |
| DeepSeek | 1 (Q4) | cites Claude, GPT |
| Grok | 0 times | cites Claude x4 |
The Grok case deserves attention. Asked four times, including «the best model» where nothing stopped it from citing itself, it never mentioned Grok. It recommended Claude every time. At the other extreme, Qwen places itself first on the first two questions (Qwen-Max, then Qwen-Coder) before yielding on the last two.
Takeaway #3: the temporal divide, the real revealer
Here's the angle no one looks at, yet the most instructive. Not all AIs share the same knowledge cutoff. Some know the mid-2026 market; others answer as if it were still 2024. Their responses betray this divide.
Freshness of market knowledge
| 🟢Up to date (mid-2026 models) | 🔴Outdated (2024 models) | |
|---|---|---|
| Claude | Fable 5, Opus 4.8, Gemini 3.1 | — |
| Gemini | Fable 5, GPT-5.5, DeepSeek V4 | — |
| Grok | Fable 5, Opus 4.8 | — |
| Kimi | Mythos, Opus 4.7, Gemini 3.1 | — |
| ChatGPT | GPT-5, Opus 4.1 | Opus 4.1 already dated |
| Qwen | — | Claude 3.5 Sonnet |
| Mistral | — | GPT-4o, Claude 3.5 |
| DeepSeek | — | OpenAI o1, Claude 3.7 |
Mistral recommends «GPT-4o» as the best model overall — a 2024 model. DeepSeek cites «OpenAI o1» and «Claude 3.7 Sonnet», also largely outdated by July 2026. Qwen mentions «Claude 3.5 Sonnet» for writing. These answers aren't wrong out of malice: they simply reflect an older training date, or a less frequently updated knowledge base.
What this study changes in how you use AI
Beyond the anecdote, three concrete reflexes to adopt, whatever AI you use.
First, neutralize self-preference. Never ask a model if it's the best for your task; ask which other model it would recommend. You'll get a far more reliable opinion.
Second, date the knowledge. Before following a tool, version or price recommendation from an AI, ask: does this model know what exists today? When in doubt, cross-check with a current source. A confident answer isn't a recent one.
Finally, read the consensus, not the isolated opinion. A single model can be wrong or outdated. But when several independent AIs converge on the same name, the signal becomes usable. That's exactly the principle of a comparator: aggregate rather than trust a single source.
🎯 Compare models on up-to-date data
Rather than trusting an AI stuck in 2024: 30+ models compared in real time — pricing, benchmarks, language, sovereignty. Verified and sourced.
Test your understanding
What's the best way to get an honest opinion from an AI about the best model?
Methodology and limits
Study conducted in July 2026, eight AIs queried via their free consumer interfaces, once each, with the same four questions phrased identically. Responses reflect each model's first spontaneous output, without follow-up.
Limits to keep in mind: a model's answers vary run to run (non-zero temperature); free interfaces don't always give access to a model's latest version, which can amplify the freshness effect; and the panel, though representative of the major labs, isn't exhaustive. This study doesn't establish a quality ranking of models: it observes what models say about each other, which is a different — and complementary — datapoint to objective benchmarks.
For a ranking based on measurements rather than opinions, see our comparator and the monthly AI Podium, updated from verified sources.