LIVE
Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|
🇨🇳 China🧠 GeneralistReleased December 2024

DeepSeek V3

DeepSeek
Open-sourceFreeSelf-hostRGPD
Verdict

A solid pick for self-hosting and full data control.

Overview

DeepSeek V3 is a revolution: 671B parameters in MoE, performance close to GPT-4o, but 30x lower price and MIT license (commercial use, self-hostable). Limits: French performance slightly behind, and for European enterprise, the official API is in China (prefer self-hosted deployment).

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1396
Verified benchmarks
GPQA59%
SWE-Bench42%
MATH90.2%
MMLU87.5%
HumanEval89%

What do these scores mean?

GPQA59%solid

PhD-level science questions (physics, chemistry, biology), with no tool access.

SWE-Bench42%fair

Resolving real GitHub issues under real conditions (SWE-bench Verified).

MATH90.2%world-class

Competition-level math problems.

MMLU87.5%excellent

General knowledge across dozens of academic subjects.

HumanEval89%excellent

Python code generation from specifications.

Strengths

  • Excellent at code
  • World-class on Arena
  • Has a free tier
  • Open-source and self-hostable

Limitations

  • Average speed

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you have GDPR constraints
Skip it if…
  • you need very fast responses

Ideal use cases

  • Ultra-low-cost coding
  • Self-hosting
  • Academic research
  • Budget-tight apps

Access & availability

Paid API per tokenFree tier availableSelf-hostable (open weights)

Key specifications

Context
128K
Input price
$0.27/M
Output price
$1.10/M
Speed
60 tok/s
Price not auditedBenchmarks not audited

Estimate your monthly cost

Per-token API pricing
$16/ month with DeepSeek V3

For the same usage

Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

Privacy

GDPR-compliantHosting : CN/selfAnonymizable data

Run it locally

Dedicated GPU server1.1M3k
QuantizationDiskRAM / VRAMTypical hardware
Q4 · recommended390.4 GB431 GBServer GPU infra (H100/A100…)
Q8 · balanced733 GB808 GBServer GPU infra (H100/A100…)
FP16 · max quality1370 GB1509 GBServer GPU infra (H100/A100…)

Estimates for a moderate context. Long contexts need more RAM (KV cache).

🤗 Hugging Face
ollama run deepseek-v3

Deploy

Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.

Estimated memory · FP16 1509 Go · Q4 431 Go

Easiest way to try it on a workstation. Install Ollama, then:

ollama run deepseek-v3
deepseek-ai/DeepSeek-V3-0324
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
MoE
Size
671B (37B actifs MoE)
Cutoff date
Jun 2024
Inputs
text
Outputs
text, code
License
mit
Hosting
CN/self
Hallucination score
3/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0019
1M in + 1M out
≈ $1.37
All benchmarks
Arena Elo1396
MMLU87.5%
GPQA59.0%
HumanEval89.0%
SWE-Bench42.0%
MATH90.2%

Similar models

Articles mentioning it