LIVE
Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|Introducing Cosmos 3 Edge20/07/26 · Hugging Face|Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling20/07/26 · Anthropic|At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI20/07/26 · NVIDIA|China's open-weights AI strategy is winning20/07/26|Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin20/07/26 · NVIDIA|Safety and alignment in an era of long-horizon models20/07/26 · OpenAI|Claude Fable produced a counterexample to the Jacobian Conjecture20/07/26 · Anthropic|Apply for Anthropic’s AI for Science rare disease research grants20/07/26 · Anthropic|AI advice made people less accurate but more confident – sudy19/07/26|Vincentwei1021/video-shotcraft: AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 106 shot recipe cards, 161 motion previews, a production-ready template19/07/26 · Anthropic|Claude Code uses Bun written in Rust now19/07/26 · Anthropic|OpenAI reduces Codex Model Context Size from 372k to 272k19/07/26 · OpenAI|
🇺🇸 United States🧠 GeneralistReleased April 2025

Llama 4 Scout

Meta
Open-sourceFreeSelf-host
Verdict

A solid pick for self-hosting and full data control.

Overview

Llama 4 Scout holds the absolute context record (10 million tokens) on a deployable-on-one-GPU model. Ideal for analyzing entire document bases. Its MoE version activates only 17B params, fast even on accessible hardware.

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1290
Verified benchmarks
MMLU80%
HumanEval84%

What do these scores mean?

MMLU80%excellent

General knowledge across dozens of academic subjects.

HumanEval84%excellent

Python code generation from specifications.

Strengths

  • Strong French quality
  • Very long context
  • Has a free tier
  • Open-source and self-hostable

Limitations

  • No native GDPR guarantee

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you work in French
Skip it if…
  • you handle sensitive EU data

Ideal use cases

  • Massive documents (millions of tokens)
  • Modest self-host
  • Research
  • Mid-hardware apps

Access & availability

Paid API per tokenFree tier availableSelf-hostable (open weights)

Key specifications

Context
10M
Input price
$0.40 $/M
Output price
$1.20 $/M
Speed
90 tok/s
Price not auditedBenchmarks not audited

Estimate your monthly cost

Per-token API pricing
$21/ month with Llama 4 Scout

For the same usage

Indicative estimate based on standard API rates (excluding caching, batch and volume discounts). Always check the official pricing before committing.

Privacy

Not GDPR-compliantHosting : US/selfAnonymizable data

Run it locally

Dedicated GPU server722k1k
QuantizationDiskRAM / VRAMTypical hardware
Q4 · recommended62.1 GB70 GB96-128GB Mac Studio / 4× GPUs
Q8 · balanced116.6 GB130 GBServer GPU infra (H100/A100…)
FP16 · max quality218 GB242 GBServer GPU infra (H100/A100…)

Estimates for a moderate context. Long contexts need more RAM (KV cache).

🤗 Hugging Face
ollama run llama4:scout
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
MoE
Size
109B (17B actifs MoE)
Cutoff date
Nov 2024
Inputs
text, image
Outputs
text, code
License
llama-community
Hosting
US/self
Hallucination score
3/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
≈ $0.0024
1M in + 1M out
≈ $1.60
All benchmarks
Arena Elo1290
MMLU80.0%
GPQA
HumanEval84.0%
SWE-Bench
MATH

Similar models