LIVE
US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|
Beginner🤖

GPT-5: everything it really does in 2026

GPT-5 released August 2025: 71% on SWE-Bench (vs 50% GPT-4), integrated Sora, autonomous agents. We tested everything for 6 months, strengths, limits, prices.

14 min readPublished May 6, 2026· Updated September 17, 2026

In one sentence

GPT-5 was released in August 2025 and redefined what we expect from an AI chatbot. More powerful, faster, more creative than GPT-4, it natively integrates Sora (video), DALL-E 3 (image), Whisper (voice), Canvas (collaborative editing), and Agents (autonomous actions). But behind the marketing, what does it actually do? We tested everything.

🤖
The analogy that works
GPT-4 vs GPT-5 is like Excel 2016 vs Microsoft 365. Excel 2016 does 80% of the job, GPT-5 does 95% : and with a smoother interface, fewer errors, and an integrated suite (Sora, DALL-E, Canvas, agents). The difference isn't spectacular line by line, but accumulated over 1 year of use, it's a different era.

🥊 Want to compare GPT-5 with Claude and Gemini?

We've made a detailed comparison of the 3 dominant AI chatbots in 2026.

See the comparison

The GPT timeline: where are we?

The evolution of GPT (OpenAI)

  1. GPT-1

    First model (117M parameters). Proves it works, but unusable in practice.

  2. GPT-2

    1.5B parameters. First fears: 'AI will generate fake news'. OpenAI hesitates to release it.

  3. GPT-3

    175B parameters. First shock: we can really converse. Birth of prompt engineering.

  4. GPT-3.5 (ChatGPT)

    The general public discovers AI. 100M users in 2 months. Global revolution.

  5. GPT-4

    Major leap in reasoning. Multimodal (image). Becomes usable for professional work.

  6. GPT-4o

    Native multimodal (real-time voice). Natural voice mode.

  7. o1, o3 (reasoning)

    Models specialized in reasoning. Math, code, sciences.

  8. GPT-5

    Unification of all capabilities. Integrated agents. Video (Sora). The 2026 standard.

  9. GPT-5.5?

    Incremental improvements expected. No major leap before 2027.

What GPT-5 does BETTER than GPT-4

GPT-5 vs GPT-4 performance on 2026 benchmarks
📊 GPT-4 (grey) vs GPT-5 (cyan) on 8 benchmarks 0% 25% 50% 75% 100% Knowledge (MMLU) 86% 89% Code (SWE-Bench) 50% 71% Math (AIME) 13% 85% Sciences (GPQA) 39% 77% French (quality) 74% 82% Factuality (low hallu.) 61% 79% Speed (tokens/sec) 36 t/s 87 t/s Context (tokens) 128K 400K GPT-4 GPT-5
On 8 tested categories, GPT-5 surpasses GPT-4 everywhere. The gaps are notable.

The 8 flagship capabilities of GPT-5

📚Complete tour of GPT-5 features in 2026

1. 🧠 Multi-step reasoning (Reasoning Mode)

GPT-5 inherited the o1/o3 system: it thinks before answering. For complex problems (math, code, strategy), it generates an internal chain-of-thought.

Use case: "Design me a pricing strategy for my B2B SaaS in 5 steps." → GPT-5 will analyze, compare, justify, not just spit out text.

Limitation: 2-10x slower than normal mode. Reserved for tasks that deserve it.

2. 💻 Code (truly) production-ready

71% on SWE-Bench (= it solves 71% of real bugs in open-source projects). That's huge vs 50% for GPT-4.

Concrete use cases:

  • Generate a full-stack app (Next.js + Postgres + auth) in 1 prompt
  • Multi-file refactoring while maintaining consistency
  • Debug complex bugs with stack trace

Limitation: still inferior to Claude Opus 4.7 (78%) on this benchmark.

3. 🎬 AI Video (Sora integrated)

ChatGPT Plus includes limited Sora (60 videos/month). ChatGPT Pro = unlimited.

Use case: "Generate me a 10s video of a woman drinking coffee in Paris at sunset, close-up shot." → Video in 5 min.

Limitation: 60 seconds max, hands sometimes weird, no sound (must add with Whisper or ElevenLabs).

4. 🎨 Native images (DALL-E 3)

Image generation integrated into the conversation. You can edit an image (change a detail, the style).

Use case: rapid prototyping, blog illustrations, design mockups.

Limitation: not as good as Midjourney v8 on photorealistic quality, but more convenient because integrated.

5. 🎙️ Voice mode (real conversation)

You speak, GPT-5 responds in real-time, with natural voice, intonations. You can interrupt, it resumes. Like talking to a human.

Use case: driving assistant, brainstorming while walking, language learning.

Limitation: it hears you too well sometimes (interrupts if you hesitate too much).

6. 📝 Canvas (collaborative editing)

Instead of having a wall of text, Canvas gives you an editable document where the AI and you collaborate. Select a sentence → "rewrite this shorter" → targeted modification.

Use case: writing long articles, code review, legal contracts.

Killer feature: you can request inline suggestions (like Google Docs).

7. 🤖 Agents (autonomous actions)

GPT-5 can execute actions: browse the web, fill forms, make purchases, send emails (with your approval).

Use case: "Book me an Uber for 7 PM to Gare du Nord." → the agent opens Uber, configures the booking, asks for your validation.

Limitation: still slow (1-5 min for a simple task), sometimes blocked by CAPTCHAs, risky on financial actions.

8. 📊 Custom GPTs (specialized assistants)

You create your own GPT with custom instructions, personal files, APIs. 1M+ public Custom GPTs.

Examples:

  • "Accounting GPT" trained on your accounting
  • "Marketing GPT" with your brand
  • "Recipe GPT" with your food preferences

What GPT-5 does NOT do well

Don't be naive: the real limitations
1. Hallucinations still present -40% vs GPT-4, but not zero. On specific topics (case law, medicine, precise dates), it sometimes invents. → Always verify important facts. 2. Dated knowledge Its knowledge base stops at April 2026 (at time of writing). For news, you need to activate web search (integrated but limited to 50 searches/month in Plus). 3. Bad at practical physics Ask it to virtually juggle 3 balls: it fails. Understanding of the real physical world = limited. (Good news: Gemini, with Genie 3, is better on this aspect.) 4. Bad with very precise numbers Mental calculation: OK. Calculation to 12 decimals: to be verified. Always use a real calculator for critical numbers. 5. Not really "intelligent" GPT-5 is a sophisticated autocomplete. It predicts the continuation of your text, doesn't truly understand. On deep philosophical questions, it shows. 6. Costly in API If you make an app that calls the GPT-5 API 100K times/month, you can quickly pay £1,000-5,000/month. Compare with Claude (often cheaper) or open-source models.

How to use it like a PRO

🎯
'TAGE' method for optimized prompts
T : Task: be precise about what you want ❌ "Write an email" → ✅ "Write a commercial follow-up email to a B2B SaaS prospect" A : Audience: for whom? ❌ (nothing) → ✅ "The prospect is CFO of a 50-person SME, technical, short on time" G : Guidelines: style, format, length ❌ (nothing) → ✅ "Tone: direct, not salesy. Format: 3 paragraphs max. Include 1 open question." E : Example: show what you expect ❌ (nothing) → ✅ "Here's a similar email I really liked: [example]" Bonus: if answer not good enough, say "improve" → GPT-5 refines.

GPT-5 vs competitors: which to choose?

GPT-5 facing the competition

 🥊Criterion🏆The best
General versatilityGPT-5 ★★★★★ / Claude ★★★★★ / Gemini ★★★★GPT-5 (slight lead)
Complex codeGPT-5 71% / Claude 78% / Gemini 65%Claude Opus 4.7
Writing / French styleGPT-5 ★★★ / Claude ★★★★★ / Gemini ★★★★Claude
Multimodal (video, image, voice)GPT-5 ★★★★★ / Claude ★★ / Gemini ★★★★★GPT-5 (Sora integrated)
Long context (entire book)GPT-5 400K / Claude 500K / Gemini 1MGemini (1M tokens)
Standard monthly priceGPT-5 Plus £20 / Claude Pro £20 / Gemini Pro £23Tie
Ecosystem (plugins, GPTs)ChatGPT massive / Claude limited / Gemini ★★★GPT-5 / ChatGPT
Voice modeGPT-5 ★★★★★ (advanced voice mode)GPT-5

TL;DR:

  • GPT-5 = the all-rounder with the best ecosystem
  • Claude = for writing and coding
  • Gemini = for very long contexts and Google integration

→ The real pro uses all 3 depending on tasks.

Should you get Plus (£20) or Pro (£200)?

Decision help
ChatGPT Plus (£20/month) is enough if you: - ✅ Use ChatGPT 1-3h/day - ✅ Mainly do text / basic code - ✅ Want limited Sora (60 videos/month) - ✅ You're starting or testing → 90% of users are OK with Plus ChatGPT Pro (£200/month) is worth it if you: - ✅ Are a professional dev or video content creator - ✅ Use Sora every day - ✅ Generate code in production (architect) - ✅ You make complex agents - ✅ You bill (clients) or resell what you produce → Justified as soon as your ROI > £200/month Quick calculation: if GPT-5 saves you >5h/month that Plus doesn't cover, Pro is profitable.

What comes after GPT-5

Rumours from the OpenAI 2026-2027 roadmap
OpenAI doesn't officially communicate, but according to leaks and clues: GPT-5.5 (end 2026?) - Incremental improvements (performance, speed, hallucinations) - No major leap GPT-6 (2027?) - Major leap expected: truly autonomous agents (multi-day) - Long-term planning capability - First pre-AGI model according to some O5 / Strawberry / Q* (rumours) - Even more advanced reasoning model - Capable of original mathematical discoveries → Advice: don't wait for GPT-6 to get started. GPT-5 is already incredible, it's up to you to learn it.

The metaphor that sums it all up

🚀
The shift from car to rocket
GPT-3 (2020) was like driving a 2CV: it worked, it impressed, but limited (30 mph, 4 seats). GPT-4 (2023) = the Tesla Model 3: comfortable, more powerful, but expensive and with battery limitations. GPT-5 (2025) = the reusable SpaceX rocket: you can do things impossible with previous cars (AI video, agents, native multimodal). But you must learn to pilot differently, it's not a car anymore, it's a spacecraft. → Those who learn to pilot it in 2026 will be 5-10 years ahead of those who start in 2030.

Absolutely remember

  • GPT-5 released August 2025, the 2026 standard
  • +25% reasoning, +40% code, -40% hallucinations vs GPT-4
  • Sora video, DALL-E images, Whisper voice, Canvas, Agents integrated
  • Real limitations: still hallucinations, bad at physics, costly in API
  • TAGE method for optimized prompts
  • Plus £20 enough for 90% of users; Pro £200 for intensive pros
  • To combine with Claude (code/writing) and Gemini (long context)

GPT-5 is the most used AI in the world in 2026 (700M active users). If you haven't started using it seriously yet, start this week.

🧠 Quiz
Question 1 of 3

What performance leap does GPT-5 show vs GPT-4 on the code benchmark (SWE-Bench)?

To go further

Tags
GPT-5OpenAIChatGPTCapacités IA

Read next