Grok 4
A versatile model, balanced across most use cases.
Overview
Grok 4 stands out for its live access to the X (Twitter) feed and lighter safety filters than competitors. Appreciated for real-time monitoring and less constrained conversations. French quality is OK but trailing. Avoid for European enterprise use cases (limited GDPR).
Skill profile
What do these scores mean?
PhD-level science questions (physics, chemistry, biology), with no tool access.
Resolving real GitHub issues under real conditions (SWE-bench Verified).
Competition-level math problems.
General knowledge across dozens of academic subjects.
Python code generation from specifications.
Strengths
- Excellent at code
- World-class on Arena
Limitations
- Average speed
- No native GDPR guarantee
- Light safety filters
Who is it for
- you build with a coding agent
- you handle sensitive EU data
- you need very fast responses
- you want strict guardrails
Ideal use cases
- Real-time X search
- Unfiltered conversations
- Social media analysis
- Humor
- Monitoring
Access & availability
Key specifications
Estimate your monthly cost
Per-token API pricingFor the same usage
- DeepSeek V3$16-94%
- GPT-5$108-58%
- Claude Sonnet 4.6$196-23%
Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.
Privacy
Advanced data · for expertsArchitecture, modalities, detailed cost, full benchmarks▾
| Arena Elo | 1360 |
| MMLU | 87.0% |
| GPQA | 72.0% |
| HumanEval | 88.5% |
| SWE-Bench | 60.0% |
| MATH | 87.5% |