LIVE
Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|
🇨🇳 China generalistReleased July 2026

Kimi K3

Moonshot AI
Open-sourceSelf-host
Verdict

A solid pick for self-hosting and full data control.

Overview

Kimi K3, from Moonshot AI, is the first open 3T-class model: 2.8 trillion parameters (104B activated), a 896-expert MoE architecture, and a 1-million-token context with native text/image/video understanding. It targets long-horizon agentic work — coding, research, office automation. Real limit: its raw knowledge score (HLE-Full) trails Claude Fable 5. Weights are open under the Kimi K3 licence and can run on European GPU servers, but its scale and Chinese origin call for careful governance.

Skill profile

Not disclosed

Strengths

  • Very long context
  • Open-source and self-hostable

Limitations

  • No native GDPR guarantee
  • API pricing not disclosed

Who is it for

A good fit if…
  • you want to control cost or self-host
Skip it if…
  • you handle sensitive EU data

Ideal use cases

  • Long-horizon autonomous coding
  • Deep research automation
  • Office task automation
  • Video and image analysis
  • Multi-tool agent orchestration

Access & availability

Self-hostable (open weights)

Key specifications

Context
1M
Input price
Output price
Speed
Price not auditedBenchmarks not audited

Privacy

Not GDPR-compliant

Run it locally

Dedicated GPU server2.6M11k
QuantizationDiskRAM / VRAMTypical hardware
Q4 · recommended1596 GB1758 GBServer GPU infra (H100/A100…)
Q8 · balanced2996 GB3298 GBServer GPU infra (H100/A100…)
FP16 · max quality5600 GB6162 GBServer GPU infra (H100/A100…)

Estimates for a moderate context. Long contexts need more RAM (KV cache).

Deploy

Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.

Estimated memory · FP16 6162 Go · Q4 1758 Go2 GPUs

OpenAI-compatible server for production on NVIDIA GPUs.

pip install vllm
vllm serve moonshotai/Kimi-K3 \
  --max-model-len 32768 \
  --tensor-parallel-size 2 \
  --dtype auto
moonshotai/Kimi-K3
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
MoE
Size
2800B (MoE, 104B actifs)
Cutoff date
Inputs
text
Outputs
text
License
kimi-k3
Hosting
Hallucination score
Estimated cost (API)
Typical exchange (~3k in / 1k out)
Not disclosed
1M in + 1M out
Not disclosed
All benchmarks
Arena Elo
MMLU
GPQA
HumanEval
SWE-Bench
MATH

Similar models