🇺🇸 United States🧠 GeneralistReleased April 2025
Llama 4 Scout
Meta
Open-sourceFreeSelf-host
Verdict
A solid pick for self-hosting and full data control.
Overview
Llama 4 Scout holds the absolute context record (10 million tokens) on a deployable-on-one-GPU model. Ideal for analyzing entire document bases. Its MoE version activates only 17B params, fast even on accessible hardware.
Skill profile
Arena Elo1290
Verified benchmarks
MMLU80%
HumanEval84%
What do these scores mean?
MMLU80%excellent
General knowledge across dozens of academic subjects.
HumanEval84%excellent
Python code generation from specifications.
Strengths
- Strong French quality
- Very long context
- Has a free tier
- Open-source and self-hostable
Limitations
- No native GDPR guarantee
Who is it for
A good fit if…
- you want to control cost or self-host
- you work in French
Skip it if…
- you handle sensitive EU data
Ideal use cases
- Massive documents (millions of tokens)
- Modest self-host
- Research
- Mid-hardware apps
Access & availability
Paid API per tokenFree tier availableSelf-hostable (open weights)
Key specifications
Context
10M
Input price
$0.40 $/M
Output price
$1.20 $/M
Speed
90 tok/s
Price not auditedBenchmarks not audited
Estimate your monthly cost
Per-token API pricing$21/ month with Llama 4 Scout
For the same usage
- GPT-5$108+428%
- Claude Sonnet 4.6$196+857%
- Grok 4$256+1150%
Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.
Privacy
Not GDPR-compliantHosting : US/selfAnonymizable data
Run it locally
Dedicated GPU server722k1k
| Quantization | Disk | RAM / VRAM | Typical hardware |
|---|---|---|---|
| Q4 · recommended | 62.1 GB | 70 GB | 96-128GB Mac Studio / 4× GPUs |
| Q8 · balanced | 116.6 GB | 130 GB | Server GPU infra (H100/A100…) |
| FP16 · max quality | 218 GB | 242 GB | Server GPU infra (H100/A100…) |
Estimates for a moderate context. Long contexts need more RAM (KV cache).
🤗 Hugging Face
ollama run llama4:scout
Advanced data · for expertsArchitecture, modalities, detailed cost, full benchmarks▾
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
MoE
Size
109B (17B actifs MoE)
Cutoff date
Nov 2024
Inputs
text, image
Outputs
text, code
License
llama-community
Hosting
US/self
Hallucination score
3/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0024
1M in + 1M out
≈ $1.60
All benchmarks
| Arena Elo | 1290 |
| MMLU | 80.0% |
| GPQA | — |
| HumanEval | 84.0% |
| SWE-Bench | — |
| MATH | — |