LIVE
Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|
≤ 2 $

Best API LLMs under $2 per million output tokens

API-accessible models with an output price at or under $2 per million tokens, ranked by Elo. The real quality-to-cost ratio for high-volume pipelines — classification, extraction, summarisation — where price per token matters more than the last benchmark point.

#ModelArena EloOutput / M tokensHostingLicence
1🇨🇳GLM-5
Zhipu AI
1458$2.00/MSelf-hostableapache-2.0
2🇨🇳Qwen3.7 Plus
Alibaba
1455$1.60/MAPIproprietary
3🇨🇳DeepSeek V3.2
DeepSeek
1425$0.42/MSelf-hostablemit
4🇨🇳DeepSeek V3
DeepSeek
1396$1.10/MSelf-hostablemit
5🇺🇸GPT-5 mini
OpenAI
1389$2.00/MEU optionproprietary
6🇨🇳Yi-Lightning
01.AI
1328$0.14/MAPIproprietary
7🇺🇸Llama 4 Scout
Meta
1321$1.20/MSelf-hostablellama-4-community
8🇨🇳DeepSeek Coder V3
DeepSeek
1280$0.28/MSelf-hostablemit
9🇫🇷Codestral
Mistral AI
1240$0.90/MSelf-hostablemistral-non-commercial
10🇫🇷Mistral Small 3
Mistral AI
1235$0.60/MSelf-hostableapache-2.0
Data verified on 4 September 2026Open in the comparator

Method

Deterministic ranking on comparator data: LMArena Elo (style control) first, HuggingFace downloads second for models without an Elo. Prices and licences are verified at the official source and dated. No model is ranked by an LLM.

Other guides