LIVE
Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|
🇺🇸 United States🧠 GeneralistReleased July 2025

Grok 4

xAI
Verdict

A versatile model, balanced across most use cases.

Overview

Grok 4 stands out for its live access to the X (Twitter) feed and lighter safety filters than competitors. Appreciated for real-time monitoring and less constrained conversations. French quality is OK but trailing. Avoid for European enterprise use cases (limited GDPR).

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1360
Verified benchmarks
GPQA72%
SWE-Bench60%
MATH87.5%
MMLU87%
HumanEval88.5%

What do these scores mean?

GPQA72%very good

PhD-level science questions (physics, chemistry, biology), with no tool access.

SWE-Bench60%solid

Resolving real GitHub issues under real conditions (SWE-bench Verified).

MATH87.5%excellent

Competition-level math problems.

MMLU87%excellent

General knowledge across dozens of academic subjects.

HumanEval88.5%excellent

Python code generation from specifications.

Strengths

  • Excellent at code
  • World-class on Arena

Limitations

  • Average speed
  • No native GDPR guarantee
  • Light safety filters

Who is it for

A good fit if…
  • you build with a coding agent
Skip it if…
  • you handle sensitive EU data
  • you need very fast responses
  • you want strict guardrails

Ideal use cases

  • Real-time X search
  • Unfiltered conversations
  • Social media analysis
  • Humor
  • Monitoring

Access & availability

Paid API per tokenConsumer subscription (~30€/mo)

Key specifications

Context
256K
Input price
$5.00 $/M
Output price
$15.00 $/M
Speed
60 tok/s
Price not auditedBenchmarks not audited

Estimate your monthly cost

Per-token API pricing
$256/ month with Grok 4

For the same usage

Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

Privacy

Not GDPR-compliantHosting : US
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
MoE
Size
undisclosed
Cutoff date
Mar 2025
Inputs
text, image
Outputs
text, code
License
commercial
Hosting
US
Hallucination score
3/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0300
1M in + 1M out
≈ $20.00
All benchmarks
Arena Elo1360
MMLU87.0%
GPQA72.0%
HumanEval88.5%
SWE-Bench60.0%
MATH87.5%

Similar models