Qwen3-Coder 480B
A solid pick for self-hosting and full data control.
Overview
Qwen3-Coder-480B-A35B is Alibaba's most ambitious coding model yet: a 480-billion-parameter MoE (35B active) built to reason over entire repositories, with a native 256K-token context extendable to 1M via Yarn. It targets agentic coding and tool use, with results the vendor claims rival Claude Sonnet — a comparison that still awaits independent verification. One limit: it ships without an explicit reasoning or thinking mode. Released under Apache 2.0, it can be self-hosted on European servers to keep data local, though its scale demands a serious multi-GPU cluster.
Skill profile
Not disclosed
Strengths
- World-class on Arena
- Open-source and self-hostable
Limitations
- API pricing not disclosed
Who is it for
- you build with a coding agent
- you want to control cost or self-host
- you have GDPR constraints
Ideal use cases
- Autonomous coding agents
- Whole-repo refactoring
- Dev tool automation
- Large-scale code review
- Sovereign on-prem deployment
Access & availability
Key specifications
Privacy
Run it locally
| Quantization | Disk | RAM / VRAM | Typical hardware |
|---|---|---|---|
| Q4 · recommended | 273.6 GB | 303 GB | Server GPU infra (H100/A100…) |
| Q8 · balanced | 513.6 GB | 567 GB | Server GPU infra (H100/A100…) |
| FP16 · max quality | 960 GB | 1058 GB | Server GPU infra (H100/A100…) |
Estimates for a moderate context. Long contexts need more RAM (KV cache).
Deploy
Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.
Easiest way to try it on a workstation. Install Ollama, then:
ollama run qwen3-coder:480bAdvanced data · for expertsArchitecture, modalities, detailed cost, full benchmarks▾
| Arena Elo | 1387 |
| MMLU | — |
| GPQA | — |
| HumanEval | — |
| SWE-Bench | — |
| MATH | — |