24 Go
Best local LLM for 24 GB VRAM
Open-weight models that fit a single 24 GB card (RTX 3090/4090 class) in Q4 quantisation, ranked by LMArena Elo. This is the tier where a workstation runs a genuinely useful model for document analysis, code and RAG without any data leaving the machine.
| # | Model | Arena Elo | RAM (Q4) | Hosting | Licence |
|---|---|---|---|---|---|
| 1 | 🇺🇸Gemma 4 31B Google | 1451 | 21 Go | Self-hostable | apache-2.0 |
| 2 | 🇺🇸Gemma 4 26B-A4B Google | 1438 | 18 Go | Self-hostable | apache-2.0 |
| 3 | 🇺🇸Gemma 3 27B Google | 1365 | 19 Go | Self-hostable | gemma |
| 4 | 🇨🇳Qwen3 32B Alibaba | 1347 | 23 Go | Self-hostable | apache-2.0 |
| 5 | 🇺🇸Gemma 3 12B Google | 1342 | 9 Go | Self-hostable | gemma |
| 6 | 🇨🇳QwQ 32B Alibaba | 1336 | 23 Go | Self-hostable | apache-2.0 |
| 7 | 🇺🇸gpt-oss-20b OpenAI | 1317 | 15 Go | Self-hostable | apache-2.0 |
| 8 | 🇺🇸Gemma 3n E4B Google | 1317 | 7 Go | Self-hostable | gemma |
| 9 | 🇺🇸Nemotron 3 Nano 30B-A3B NVIDIA | 1314 | 21 Go | Self-hostable | nvidia-open-model |
| 10 | 🇺🇸Granite 4.1 8B IBM | 1306 | 7 Go | Self-hostable | apache-2.0 |
Data verified on 4 September 2026Open in the comparator
Method
Deterministic ranking on comparator data: LMArena Elo (style control) first, HuggingFace downloads second for models without an Elo. Prices and licences are verified at the official source and dated. No model is ranked by an LLM.