LIVE
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out03/09/26 · Anthropic|OpenAI's GPT-6 Astra on ARC-AGI-303/09/26 · OpenAI|GPT-6 Astra03/09/26 · OpenAI|OpenAI begins rolling out GPT-6 Astra03/09/26 · OpenAI|ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize03/09/26 · Anthropic|Sparks Fly: NVIDIA Accelerates Local AI at IFA 202603/09/26 · NVIDIA|Introducing WeatherNext 3, our most advanced and accurate global weather AI model03/09/26 · Google|Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly03/09/26 · Anthropic|Claude outage – Resolved03/09/26 · Anthropic|Daybreak for Frontline Defenders: $1B to protect essential services03/09/26 · OpenAI|NeoMME: an efficient Multimodal-native and Multilingual Encoder03/09/26 · Hugging Face|‘NBA 2K27’ With NVIDIA DLSS 5 Leads 28 New Games Coming to GeForce NOW03/09/26 · NVIDIA|
🇦🇪 UAE🧠 GeneralistReleased September 2023

Falcon 180B

TII
Open-sourceFreeSelf-host
Verdict

A solid pick for self-hosting and full data control.

Overview

Falcon 180B is developed by Abu Dhabi's Technology Innovation Institute. Heavily used in the MENA region and Arabic-language apps. Older than newer models, performance trailing, but remains one of the historic open-source pillars.

Skill profile

FrenchReasoningSpeedCreativitySafety
Arena Elo1147
Verified benchmarks
MMLU70.5%
HumanEval35%

What do these scores mean?

MMLU70.5%very good

General knowledge across dozens of academic subjects.

HumanEval35%fair

Python code generation from specifications.

Strengths

  • Has a free tier
  • Open-source and self-hostable

Limitations

  • Average speed
  • No native GDPR guarantee
  • Light safety filters

Who is it for

A good fit if…
  • you want to control cost or self-host
Skip it if…
  • you handle sensitive EU data
  • you need very fast responses
  • you want strict guardrails

Ideal use cases

  • Arabic-language apps
  • MENA sovereignty
  • Academic research

Access & availability

Free tier availableSelf-hostable (open weights)

Key specifications

Context
4K
Input price
Output price
Speed
30 tok/s
Price not auditedBenchmarks not audited

Privacy

Not GDPR-compliantHosting : AE/selfAnonymizable data
Advanced data · for experts
Architecture, modalities, detailed cost, full benchmarks
Architecture
Dense
Size
180B
Cutoff date
Mar 2023
Inputs
text
Outputs
text, code
License
apache-2
Hosting
AE/self
Hallucination score
2/5
Estimated cost (API)
Typical exchange (~3k in / 1k out)
Not disclosed
1M in + 1M out
Not disclosed
All benchmarks
Arena Elo1147
MMLU70.5%
GPQA
HumanEval35.0%
SWE-Bench
MATH

Similar models