LIVE
Mistral AI Raises €3 Billion to Push Open-Weight Models to the Frontier08/09/26 · Mistral AI|OpenAI agents hijacked a German wiki before the Hugging Face breach07/09/26 · OpenAI|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Meta launches Muse Spark 1.3, a model built for agentic coding02/09/26 · Meta|Post-trained Nemotron model outscores the top human at the International Olympiad in Informatics02/09/26 · NVIDIA|Google DeepMind launches Fairwind, a proactive cyber defense program for governments and enterprises02/09/26 · Google DeepMind|Study suggests neural networks implicitly encode symbolic structures02/09/26|LibreOffice sees download surge after pledging to stay AI-free08/09/26 · LibreOffice|Danijar Hafner is building agents designed to plan for the unexpected08/09/26|OpenAI launches $5 million grant program on AI and teen development08/09/26 · OpenAI|TradingAgents: an open-source framework simulates a trading desk with LLM agents08/09/26|Mistral AI Raises €3 Billion to Push Open-Weight Models to the Frontier08/09/26 · Mistral AI|OpenAI agents hijacked a German wiki before the Hugging Face breach07/09/26 · OpenAI|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Meta launches Muse Spark 1.3, a model built for agentic coding02/09/26 · Meta|Post-trained Nemotron model outscores the top human at the International Olympiad in Informatics02/09/26 · NVIDIA|Google DeepMind launches Fairwind, a proactive cyber defense program for governments and enterprises02/09/26 · Google DeepMind|Study suggests neural networks implicitly encode symbolic structures02/09/26|LibreOffice sees download surge after pledging to stay AI-free08/09/26 · LibreOffice|Danijar Hafner is building agents designed to plan for the unexpected08/09/26|OpenAI launches $5 million grant program on AI and teen development08/09/26 · OpenAI|TradingAgents: an open-source framework simulates a trading desk with LLM agents08/09/26|
Hardware Match

Which model for NVIDIA L40S · 48 Go?

48 GB usable for the model. 86 open-weight models fit in memory.

The margin goes to context: count ~1 GB per 8k tokens for a 7B model. A small margin means a short context.

ModelEloQuantizationRAM requiredMarginSpeedRun
Gemma 4 31BGoogle · 31B1451Q8_039 Go+9 GotightHugging Face ↗
Gemma 4 26B-A4BGoogle · 26B1438Q8_033 Go+15 GocomfortableHugging Face ↗
Gemma 3 27BGoogle · 27B1365Q8_034 Go+14 Gotight ollama run gemma3:27b
Qwen3 32BAlibaba · 32.8B1347Q8_041 Go+7 Gotight ollama run qwen3:32b
Gemma 3 12BGoogle · 12B1342FP1628 Go+20 Gocomfortable ollama run gemma3:12b
QwQ 32BAlibaba · 32.8B1336Q8_041 Go+7 Gotight ollama run qwq
Gemma 3n E4BGoogle · 8B1317FP1620 Go+28 Gocomfortable ollama run gemma3n:e4b
Llama 3.3 70BMeta · 70.6B1317Q4_K_M46 Go+2 Gotight ollama run llama3.3:70b
gpt-oss-20bOpenAI · 21B1317FP1648 Go+0 Gotight ollama run gpt-oss:20b
Nemotron 3 Nano 30B-A3BNVIDIA · 30B1314Q8_037 Go+11 GotightHugging Face ↗
Granite 4.1 8BIBM · 8B1306FP1620 Go+28 GocomfortableHugging Face ↗
Gemma 3 4BGoogle · 4B1303FP1611 Go+37 Gocomfortable ollama run gemma3:4b
Qwen2.5 72BAlibaba · 72.7B1303Q4_K_M48 Go+0 Gotight ollama run qwen2.5:72b
Llama 3.1 Nemotron 70BNVIDIA · 70.6B1299Q4_K_M46 Go+2 Gotight ollama run nemotron:70b
Qwen2.5-Coder 32BAlibaba · 32.8B1270Q8_041 Go+7 Gotight ollama run qwen2.5-coder:32b
Phi-4Microsoft · 14.7B1256FP1634 Go+14 Gotight ollama run phi4
Phi-4 ReasoningMicrosoft · 14.7B1256FP1634 Go+14 Gotight ollama run phi4-reasoning
OLMo 2 32BAllen AI · 32B1251Q8_040 Go+8 GotightHugging Face ↗
CodestralMistral AI · 22.2B1240Q8_028 Go+20 Gocomfortable ollama run codestral
Ministral 8BMistral AI · 8.02B1237FP1620 Go+28 GocomfortableHugging Face ↗
Llama 3.1 8BMeta · 8.03B1211FP1620 Go+28 Gocomfortable ollama run llama3.1:8b
Mixtral 8x7BMistral AI · 46.7B1197Q4_K_M31 Go+17 Gocomfortable ollama run mixtral:8x7b
Llama 3.2 3BMeta · 3.21B1166FP169 Go+39 Gocomfortable ollama run llama3.2:3b
Llama 3.2 1BMeta · 1.24B1111FP165 Go+43 Gocomfortable ollama run llama3.2:1b
Mistral 7BMistral AI · 7.25B1110FP1618 Go+30 Gocomfortable ollama run mistral:7b
Kokoro-82MHexgrad · 0.08BFP162 Go+46 GocomfortableHugging Face ↗
ChatterboxResemble AI · 0.5BFP163 Go+45 GocomfortableHugging Face ↗
Parakeet TDT 0.6B v3NVIDIA · 0.6BFP163 Go+45 GocomfortableHugging Face ↗
Qwen3 0.6BAlibaba · 0.6BFP163 Go+45 Gocomfortable ollama run qwen3:0.6b
TripoSRStability AI · 0.5BFP163 Go+45 GocomfortableHugging Face ↗
Gemma 3 1BGoogle · 1BFP164 Go+44 Gocomfortable ollama run gemma3:1b
Stable Fast 3DStability AI · 1BFP164 Go+44 GocomfortableHugging Face ↗
TRELLISMicrosoft · 1.2BFP165 Go+43 GocomfortableHugging Face ↗
Whisper large-v3OpenAI · 1.55BFP165 Go+43 GocomfortableHugging Face ↗
Dia 1.6BNari Labs · 1.6BFP166 Go+42 GocomfortableHugging Face ↗
Gemma 4 E2BGoogle · 2BFP166 Go+42 GocomfortableHugging Face ↗
Granite 4.1 3BIBM · 3BFP169 Go+39 GocomfortableHugging Face ↗
Hunyuan3D 2.1Tencent · 3BFP169 Go+39 GocomfortableHugging Face ↗
Ministral 3 3BMistral AI · 3BFP169 Go+39 GocomfortableHugging Face ↗
Orpheus 3BCanopy Labs · 3BFP169 Go+39 GocomfortableHugging Face ↗
SmolLM3 3BHugging Face · 3.1BFP169 Go+39 GocomfortableHugging Face ↗
Voxtral Mini 3BMistral AI · 3BFP169 Go+39 GocomfortableHugging Face ↗
Phi-4 MiniMicrosoft · 3.8BFP1610 Go+38 Gocomfortable ollama run phi4-mini
FLUX.2 [klein] 4BBlack Forest Labs · 4BFP1611 Go+37 GocomfortableHugging Face ↗
Gemma 4 E4BGoogle · 4BFP1611 Go+37 GocomfortableHugging Face ↗
Qwen3 4BAlibaba · 4.02BFP1611 Go+37 Gocomfortable ollama run qwen3:4b
TRELLIS.2Microsoft · 4BFP1611 Go+37 GocomfortableHugging Face ↗
CogVideoX 5BZhipu AI · 5BFP1613 Go+35 GocomfortableHugging Face ↗
Wan 2.2 5BAlibaba · 5BFP1613 Go+35 GocomfortableHugging Face ↗
Z-Image TurboAlibaba · 6BFP1615 Go+33 GocomfortableHugging Face ↗
Lucie 7BOpenLLM France · 6.7BFP1617 Go+31 GocomfortableHugging Face ↗
OLMo 3 7BAllen AI · 7BFP1617 Go+31 GocomfortableHugging Face ↗
Apertus 8BSwiss AI · 8BFP1620 Go+28 GocomfortableHugging Face ↗
Command R7BCohere · 8BFP1620 Go+28 Gocomfortable ollama run command-r7b
DeepSeek-R1 Distill 8BDeepSeek · 8.03BFP1620 Go+28 Gocomfortable ollama run deepseek-r1:8b
Granite 3.3 8BIBM · 8.1BFP1620 Go+28 Gocomfortable ollama run granite3.3:8b
Hermes 3 8BNous Research · 8.03BFP1620 Go+28 Gocomfortable ollama run hermes3:8b
HunyuanVideo 1.5Tencent · 8.3BFP1620 Go+28 GocomfortableHugging Face ↗
Ministral 3 8BMistral AI · 8BFP1620 Go+28 GocomfortableHugging Face ↗
Qwen3 8BAlibaba · 8.2BFP1620 Go+28 Gocomfortable ollama run qwen3:8b
CodeGemma 7BGoogle · 8.5BFP1621 Go+27 Gocomfortable ollama run codegemma
InternLM3 8BShanghai AI Lab · 8.8BFP1621 Go+27 GocomfortableHugging Face ↗
EuroLLM 9BUTTER Project · 9.15BFP1622 Go+26 GocomfortableHugging Face ↗
GLM-ImageZhipu AI · 9BFP1622 Go+26 GocomfortableHugging Face ↗
Mochi 1Genmo · 10BFP1624 Go+24 GocomfortableHugging Face ↗
Falcon 3 10BTII · 10.3BFP1625 Go+23 Gocomfortable ollama run falcon3:10b
FLUX.1 [dev]Black Forest Labs · 12BFP1628 Go+20 GocomfortableHugging Face ↗
Gemma 4 12BGoogle · 12BFP1628 Go+20 GocomfortableHugging Face ↗
LTX-2.3Lightricks · 22BQ8_028 Go+20 GocomfortableHugging Face ↗
Mistral Nemo 12BMistral AI · 12.2BFP1629 Go+19 Gocomfortable ollama run mistral-nemo
Devstral SmallMistral AI · 24BQ8_030 Go+18 Gocomfortable ollama run devstral
Magistral SmallMistral AI · 24BQ8_030 Go+18 Gocomfortable ollama run magistral
Mistral Small 3.2Mistral AI · 24BQ8_030 Go+18 Gocomfortable ollama run mistral-small3.2
Voxtral Small 24BMistral AI · 24BQ8_030 Go+18 GocomfortableHugging Face ↗
Ministral 3 14BMistral AI · 14BFP1633 Go+15 GocomfortableHugging Face ↗
Qwen3.6 27BAlibaba · 27BQ8_034 Go+14 GotightHugging Face ↗
Wan 2.2 A14BAlibaba · 27BQ8_034 Go+14 GotightHugging Face ↗
Qwen3 14BAlibaba · 14.8BFP1635 Go+13 Gotight ollama run qwen3:14b
DeepSeek-Coder-V2 LiteDeepSeek · 15.7BFP1637 Go+11 Gotight ollama run deepseek-coder-v2:16b
StarCoder 2BigCode · 16BFP1637 Go+11 Gotight ollama run starcoder2:15b
HiDream-I1HiDream · 17BFP1639 Go+9 GotightHugging Face ↗
Aya Expanse 32BCohere · 32.3BQ8_040 Go+8 Gotight ollama run aya-expanse:32b
FLUX.2 [dev]Black Forest Labs · 32BQ8_040 Go+8 GotightHugging Face ↗
DeepSeek-R1 Distill 32BDeepSeek · 32.8BQ8_041 Go+7 Gotight ollama run deepseek-r1:32b
Qwen3.6 35B-A3BAlibaba · 35BQ8_043 Go+5 GotightHugging Face ↗
Qwen-Image 2512Alibaba · 20BFP1646 Go+2 GotightHugging Face ↗