FLUX.2 [klein] 4B
A solid pick for self-hosting and full data control.
Overview
FLUX.2 [klein] 4B is a text-to-image and editing model from Black Forest Labs (Germany), built for interactive, latency-critical use: sub-second inference on a single consumer GPU with as little as 13GB VRAM. It unifies text-to-image generation and multi-reference editing in one compact, Apache 2.0 licensed model cleared for commercial use. A real limit: text rendered inside images stays inaccurate, and output quality is highly sensitive to prompt phrasing. Open weights from a European vendor make fully local, on-prem deployment straightforward, with no data leaving the machine.
Skill profile
Not disclosed
Strengths
- Open-source and self-hostable
Limitations
- API pricing not disclosed
Who is it for
- you want to control cost or self-host
- you have GDPR constraints
Ideal use cases
- Local image generation
- Fast multi-reference editing
- Edge creative prototyping
- Embedded product pipelines
- Secure on-prem deployment
Access & availability
Key specifications
Privacy
Run it locally
| Quantization | Disk | RAM / VRAM | Typical hardware |
|---|---|---|---|
| Q4 · recommended | 2.3 GB | 5 GB | Any recent PC/Mac |
| Q8 · balanced | 4.3 GB | 7 GB | 16GB PC / M1+ Mac / 8GB GPU |
| FP16 · max quality | 8 GB | 11 GB | 32GB PC / 24GB Mac / RTX 4070 Ti+ |
Estimates for a moderate context. Long contexts need more RAM (KV cache).
Deploy
Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.
OpenAI-compatible server for production on NVIDIA GPUs.
pip install vllm
vllm serve black-forest-labs/FLUX.2-klein-4B \
--max-model-len 32768 \
--tensor-parallel-size 1 \
--dtype autoAdvanced data · for expertsArchitecture, modalities, detailed cost, full benchmarks▾
| Arena Elo | — |
| MMLU | — |
| GPQA | — |
| HumanEval | — |
| SWE-Bench | — |
| MATH | — |