LIVE
Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Hcompany releases NeoMME, a compact multimodal-native encoder for retrieval03/09/26 · Hcompany|Playco cuts manual fixes by 50% using GPT-6 Astra for prototyping03/09/26 · OpenAI|Legora reviews 41 financial documents in minutes with GPT-6 Astra03/09/26 · OpenAI|Hugging Face open-sources a project training a coding model to paint watercolours03/09/26 · Hugging Face|A 350M-Parameter Model Improved in 100 GRPO Steps for Better Structured Outputs03/09/26|
24 Go

Best local LLM for 24 GB VRAM

Open-weight models that fit a single 24 GB card (RTX 3090/4090 class) in Q4 quantisation, ranked by LMArena Elo. This is the tier where a workstation runs a genuinely useful model for document analysis, code and RAG without any data leaving the machine.

#ModelArena EloRAM (Q4)HostingLicence
1🇺🇸Gemma 4 31B
Google
145121 GoSelf-hostableapache-2.0
2🇺🇸Gemma 4 26B-A4B
Google
143818 GoSelf-hostableapache-2.0
3🇺🇸Gemma 3 27B
Google
136519 GoSelf-hostablegemma
4🇨🇳Qwen3 32B
Alibaba
134723 GoSelf-hostableapache-2.0
5🇺🇸Gemma 3 12B
Google
13429 GoSelf-hostablegemma
6🇨🇳QwQ 32B
Alibaba
133623 GoSelf-hostableapache-2.0
7🇺🇸gpt-oss-20b
OpenAI
131715 GoSelf-hostableapache-2.0
8🇺🇸Gemma 3n E4B
Google
13177 GoSelf-hostablegemma
9🇺🇸Nemotron 3 Nano 30B-A3B
NVIDIA
131421 GoSelf-hostablenvidia-open-model
10🇺🇸Granite 4.1 8B
IBM
13067 GoSelf-hostableapache-2.0
Data verified on 4 September 2026Open in the comparator

Method

Deterministic ranking on comparator data: LMArena Elo (style control) first, HuggingFace downloads second for models without an Elo. Prices and licences are verified at the official source and dated. No model is ranked by an LLM.

Other guides