LIVE
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|
🇺🇸 United States🔬 ReasoningReleased April 2025

o3

OpenAI
RGPD
Verdict

A good performance/compliance balance for a European company.

Overview

o3 uses a chain-of-thought architecture reinforced by RL: it "thinks" for several minutes before answering, making it extraordinarily strong on complex problems. It's the reference model for research, mathematical proofs, and algorithmic competitions. Cost and slowness reserve it for high-end use cases.

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1431
Verified benchmarks
GPQA87.7%
SWE-Bench71.7%
MATH96.7%
MMLU88.5%
HumanEval96%

What do these scores mean?

GPQA87.7%excellent

PhD-level science questions (physics, chemistry, biology), with no tool access.

SWE-Bench71.7%very good

Resolving real GitHub issues under real conditions (SWE-bench Verified).

MATH96.7%world-class

Competition-level math problems.

MMLU88.5%excellent

General knowledge across dozens of academic subjects.

HumanEval96%world-class

Python code generation from specifications.

Strengths

  • Top-tier reasoning
  • Excellent at code
  • Strong French quality
  • World-class on Arena

Limitations

  • Average speed

Who is it for

A good fit if…
  • you need advanced reasoning or analysis
  • you build with a coding agent
  • you work in French
  • you have GDPR constraints
Skip it if…
  • you need very fast responses

Ideal use cases

  • Advanced math
  • Scientific research
  • Complex algorithms
  • Proofs
  • Competitions

Access & availability

Paid API per tokenConsumer subscription (~200€/mo)

Key specifications

Context
200K
Input price
$2.00 $/M
Output price
$8.00 $/M
Speed
30 tok/s
Price not auditedBenchmarks not audited

Estimate your monthly cost

Per-token API pricing
$117/ month with o3

For the same usage

Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

Privacy

GDPR-compliantHosting : US/EUAnonymizable data
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
Dense + RL
Size
undisclosed
Cutoff date
Nov 2024
Inputs
text, image
Outputs
text, code
License
commercial
Hosting
US/EU
Hallucination score
4/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0140
1M in + 1M out
≈ $10.00
All benchmarks
Arena Elo1431
MMLU88.5%
GPQA87.7%
HumanEval96.0%
SWE-Bench71.7%
MATH96.7%

Similar models