LIVE
Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|
🇺🇸 United States🧠 GeneralistReleased February 2026

Gemini 3.1 Pro

Google
Free
Verdict

A versatile model, balanced across most use cases.

Overview

Gemini 3.1 Pro is Google DeepMind's flagship, built on a Mixture-of-Experts architecture. It leads most reasoning benchmarks (94.3% GPQA Diamond) while keeping Gemini 3 Pro's pricing ($2/$12 per million). Natively multimodal (text, image, audio, video) with a 1M-token context, it's one of the best cost/performance picks at the frontier.

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1487
Verified benchmarks
GPQA94.3%
SWE-Bench80.6%
MATH95.1%

What do these scores mean?

GPQA94.3%world-class

PhD-level science questions (physics, chemistry, biology), with no tool access.

SWE-Bench80.6%excellent

Resolving real GitHub issues under real conditions (SWE-bench Verified).

MATH95.1%world-class

Competition-level math problems.

Strengths

  • Top-tier reasoning
  • Excellent at code
  • Strong French quality
  • World-class on Arena

Limitations

  • No native GDPR guarantee

Who is it for

A good fit if…
  • you need advanced reasoning or analysis
  • you build with a coding agent
  • you work in French
Skip it if…
  • you handle sensitive EU data

Ideal use cases

  • Complex reasoning and research
  • Multimodal analysis (image, audio, video)
  • Very long context (codebases, large corpora)

Access & availability

Paid API per tokenFree tier available

Key specifications

Context
1M
Input price
$2.00 $/M
Output price
$12.00 $/M
Speed
144 tok/s
Price verified Jun 1, 2026·Official sourceBenchmarks not audited

Estimate your monthly cost

Per-token API pricing
$145/ month with Gemini 3.1 Pro

For the same usage

Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

Privacy

Not GDPR-compliant
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
Mixture-of-Experts
Size
Cutoff date
Dec 2024
Inputs
text, image, audio, video
Outputs
text
License
proprietary
Hosting
Hallucination score
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0180
1M in + 1M out
≈ $14.00
All benchmarks
Arena Elo1487
MMLU
GPQA94.3%
HumanEval
SWE-Bench80.6%
MATH95.1%

Similar models