LIVE
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|
🇫🇷 France🧠 GeneralistReleased December 2025

Mistral Large 3

Mistral AI
Open-sourceFreeSelf-hostRGPD
Verdict

A solid pick for self-hosting and full data control.

Overview

Mistral Large 3 is THE reference model for sovereign European use: France-hosted, full GDPR compliance, French team. French quality is excellent (Claude-level). Solid coding and reasoning performance, slightly behind the priciest US frontiers but at a much lower price.

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1415
Verified benchmarks
GPQA64%
SWE-Bench50%
MATH84%
MMLU84.5%
HumanEval88%

What do these scores mean?

GPQA64%solid

PhD-level science questions (physics, chemistry, biology), with no tool access.

SWE-Bench50%solid

Resolving real GitHub issues under real conditions (SWE-bench Verified).

MATH84%excellent

Competition-level math problems.

MMLU84.5%excellent

General knowledge across dozens of academic subjects.

HumanEval88%excellent

Python code generation from specifications.

Strengths

  • Excellent at code
  • Strong French quality
  • World-class on Arena
  • Has a free tier

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you work in French
  • you have GDPR constraints
Skip it if…

    Ideal use cases

    • GDPR-compliant European apps
    • European multilingual
    • French apps
    • EU enterprise

    Access & availability

    Paid API per tokenFree tier availableSelf-hostable (open weights)Consumer subscription (~15€/mo)

    Key specifications

    Context
    256K
    Input price
    $2.00 $/M
    Output price
    $6.00 $/M
    Speed
    90 tok/s
    Price not auditedBenchmarks not audited

    Estimate your monthly cost

    Per-token API pricing
    $103/ month with Mistral Large 3

    For the same usage

    Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

    Privacy

    GDPR-compliantHosting : EUAnonymizable data

    Run it locally

    Dedicated GPU server1k240
    QuantizationDiskRAM / VRAMTypical hardware
    Q4 · recommended384.7 GB425 GBServer GPU infra (H100/A100…)
    Q8 · balanced722.2 GB796 GBServer GPU infra (H100/A100…)
    FP16 · max quality1350 GB1487 GBServer GPU infra (H100/A100…)

    Estimates for a moderate context. Long contexts need more RAM (KV cache).

    Advanced data · for experts
    Architecture, modalities, detailed cost, full benchmarks
    Architecture
    Dense
    Size
    ~200B
    Cutoff date
    Feb 2025
    Inputs
    text, image
    Outputs
    text, code
    License
    commercial
    Hosting
    EU
    Hallucination score
    4/5
    Estimated cost (API)
    Typical exchange (~3k in / 1k out)
    ≈ $0.0120
    1M in + 1M out
    ≈ $8.00
    All benchmarks
    Arena Elo1415
    MMLU84.5%
    GPQA64.0%
    HumanEval88.0%
    SWE-Bench50.0%
    MATH84.0%

    Similar models