o3
A good performance/compliance balance for a European company.
Overview
o3 uses a chain-of-thought architecture reinforced by RL: it "thinks" for several minutes before answering, making it extraordinarily strong on complex problems. It's the reference model for research, mathematical proofs, and algorithmic competitions. Cost and slowness reserve it for high-end use cases.
Skill profile
What do these scores mean?
PhD-level science questions (physics, chemistry, biology), with no tool access.
Resolving real GitHub issues under real conditions (SWE-bench Verified).
Competition-level math problems.
General knowledge across dozens of academic subjects.
Python code generation from specifications.
Strengths
- Top-tier reasoning
- Excellent at code
- Strong French quality
- World-class on Arena
Limitations
- Average speed
Who is it for
- you need advanced reasoning or analysis
- you build with a coding agent
- you work in French
- you have GDPR constraints
- you need very fast responses
Ideal use cases
- Advanced math
- Scientific research
- Complex algorithms
- Proofs
- Competitions
Access & availability
Key specifications
Estimate your monthly cost
Per-token API pricingFor the same usage
- DeepSeek R2$32-73%
- Claude Opus 4.8$327+180%
- Claude Opus 4.7$327+180%
Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.
Privacy
Advanced data · for expertsArchitecture, modalities, detailed cost, full benchmarks▾
| Arena Elo | 1410 |
| MMLU | 88.5% |
| GPQA | 87.7% |
| HumanEval | 96.0% |
| SWE-Bench | 71.7% |
| MATH | 96.7% |