EN DIRECT
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|
🇺🇸 États-Unis🧠 GénéralisteSortie avril 2025

Llama 4 Scout

Meta
Open-sourceGratuitSelf-host
Verdict

Une option solide pour l'auto-hébergement et le contrôle total des données.

Présentation

Llama 4 Scout est le record absolu du contexte (10 millions de tokens) sur un modèle déployable sur un GPU. Idéal pour analyser des bases entières de documents. Sa version MoE active 17B params seulement, donc rapide même sur du hardware accessible.

Profil de compétences

FrançaisRaisonn.VitesseCréativitéSûreté
Arena Elo1323
Benchmarks vérifiés
MMLU80%
HumanEval84%

Que signifient ces scores ?

MMLU80%excellent

Connaissances générales réparties sur des dizaines de matières académiques.

HumanEval84%excellent

Génération de code Python à partir de spécifications.

Forces

  • Très bonne qualité en français
  • Très long contexte
  • Dispose d'un tier gratuit
  • Open-source et auto-hébergeable

Limites

  • Pas de garantie RGPD native

À qui ça s'adresse

Pour vous si…
  • vous voulez maîtriser vos coûts ou auto-héberger
  • vous travaillez en français
À éviter si…
  • vous manipulez des données sensibles européennes

Cas d'usage idéaux

  • Documents géants (millions de tokens)
  • Self-host modeste
  • Recherche
  • Apps avec contraintes hardware moyennes

Accès & disponibilité

API payante au tokenTier gratuit disponibleAuto-hébergeable (poids ouverts)

Caractéristiques clés

Contexte
10M
Prix entrée
$0.40 $/M
Prix sortie
$1.20 $/M
Vitesse
90 tok/s
Prix non auditéBenchmarks non audités

Estimez votre coût mensuel

Tarifs API au token
$21/ mois avec Llama 4 Scout

Pour le même usage

Estimation indicative sur les tarifs API standard (hors cache, batch et remises volume). Vérifiez toujours la grille officielle avant de vous engager.

Confidentialité

Non conforme RGPDHébergement : US/selfDonnées anonymisables

Faire tourner en local

Serveur GPU dédié499k1k
QuantizationDisqueRAM / VRAMMatériel type
Q4 · recommandé62.1 Go70 GoMac Studio 96-128 Go / 4× GPU
Q8 · équilibré116.6 Go130 GoInfra GPU serveur (H100/A100…)
FP16 · qualité max218 Go242 GoInfra GPU serveur (H100/A100…)

Estimations pour un contexte modéré. Un long contexte augmente la RAM nécessaire (KV cache).

🤗 Hugging Face
ollama run llama4:scout
Données avancées · pour experts
Architecture, modalités, coût détaillé, benchmarks complets
Architecture
MoE
Taille
109B (17B actifs MoE)
Date de coupe
nov. 2024
Entrées
text, image
Sorties
text, code
Licence
llama-community
Hébergement
US/self
Score d'hallucination
3/5
Coût estimé (API)
Échange type (~3k entrée / 1k sortie)
≈ $0.0024
1M entrée + 1M sortie
≈ $1.60
Tous les benchmarks
Arena Elo1323
MMLU80.0%
GPQA
HumanEval84.0%
SWE-Bench
MATH

Modèles similaires