Falcon Benchmark
Frontier-class Indic AI at 1/13th the serving cost.
A 26B-parameter sparse mixture-of-experts, about 4B active per query, FP4-quantized and benchmarked exactly as it shipped on a single GPU. It outscores every same-class open model measured here across 11 Indian languages, tracks models many times its size on English reasoning, and serves at $0.42 per million tokens.
- One pinned pipeline
- Greedy decoding
- Fixed seeds
- Full per-sample audit trail
- Every number re-runnable
per 1M output tokens
13× cheaper than Llama-70B, 5× than Gemma-3-27B, measured on the same GPU
MMLU-Pro
within 3 pts of DeepSeek-V3.1’s published 84.8, at about 4B active params
MILU Indic composite (gen-CoT)
best in panel across 10 Indian languages
tokens/sec @ 30 ms TTFT
12 concurrent streams, production serving stack
Serving economics: measured, identical hardware
| Benchmark | Falcon | Qwen3-30B-A3B | Gemma-3-27B | Mistral-Small-3.2 | Sarvam-M | Llama-4-Scout 109B | Llama-3.3-70B |
|---|---|---|---|---|---|---|---|
| TTFT ms @ 12512-tok prompt · lower better | 30.00 | 38.00 | 48.00 | 40.00 | 50.00 | 43.00 | 59.00 |
| Decode tok/s12 concurrent | 2,029 | 1,699 | 388 | 480 | 479 | 921 | 152 |
| $ per 1M output tokenslower better | 0.42 | 0.50 | 2.17 | 1.75 | 1.76 | 0.91 | 5.54 |
Measured on one RTX PRO 6000 (Modal $3.03/h), identical load.
Indic languages: composite scores
| Benchmark | Falcon | Qwen3-30B-A3B |
|---|---|---|
| MILU (generative CoT)avg 10 langs, India-centric MCQ | 83.2 | 73.8 |
| MMLU-ProX hard MCQavg hi/bn/mr/te, CoT | 72.9 | 55.4 |
| GSM8K-Indic mathavg 10 langs, CoT | 88.3 | 85.4 |
| GSM8K romanized HindiHinglish stressor | 86.0 | 77.2 |
| Translation en→Indicavg 10, chrF++ | 50.1 | 41.2 |
| Translation Indic→enavg 10, chrF++ | 61.6 | 58.0 |
| Hinglish→EnglishCOMI-LINGUA, 3 refs, chrF++ | 79.3 | 76.0 |
Averages over the product’s Indian languages (10 scripts + romanized Hindi). MCQ composites in %, translation in chrF++. Per-language breakdowns: results_flat.csv.
Global languages: Falcon composite scores (13 locales)
Falcon, measured| Benchmark | Falcon |
|---|---|
| Belebele reading comp.avg 13 global langs | 56.1 |
| MMLU-ProX hard MCQavg 9 langs, CoT | 75.7 |
zh, ja, ko, es, de, ru, sv, pl, it, nl, ro, ar, id. Falcon’s own measured scores across the 13 non-Indic locales the product supports; no competitor is run on this suite, so no row is marked best.
English core
| Benchmark | Falcon | Qwen3-30B-A3B | Mistral-Small-3.2 |
|---|---|---|---|
| MMLU-Pro5-shot CoT | 82.0 | 77.9 | 63.2 |
| GSM8K8-shot CoT | 94.0 | 92.1 | 83.9 |
| MATH-500math-verify | 95.2 | 96.2 | 87.6 |
| AIME 2025flex | 76.7 | 70.0 | 36.7 |
| IFEvalprompt-strict | 89.8 | 82.6 | 74.5 |
| HumanEvalpass@1 | 99.4 | 90.9 | 88.4 |
| MBPP+pass@1 | 96.6 | 92.9 | 84.9 |
All cells measured by this pipeline.
Scale contrast: bigger Meta models on the same GPU
Serving rows measured on identical hardware| Benchmark | Llama-4-Scout 109B | Llama-3.3-70B | Falcon |
|---|---|---|---|
| Total params (B) | 109.0 | 70.6 | 26.0 |
| Active params (B) | 17.0 | 70.6 | 4.0 |
| Decode tok/smeasured | 920.8 | 151.9 | 2,028.9 |
| $ / 1M tokensmeasured | 0.9 | 5.5 | 0.4 |
Serving rows measured by this pipeline on the same GPU. This table makes a cost argument only: 4.2× the parameters at 2.2× the cost (Scout) or 13× the cost (70B). Quality is deliberately not compared here: the full quality suites for these two models were not run, and setting our measured scores against vendor-published numbers taken under different protocols would not be a like-for-like comparison.
Quality per dollar
MMLU-Pro vs $/1M tokens · log scaleMMLU-Pro accuracy vs measured serving cost (log scale). Solid = both axes measured by this pipeline; ° = vendor-reported quality with our measured cost. Falcon sits at the top-left of the plotted set: the highest measured accuracy at the lowest measured cost among these models. Frontier APIs (GPT-5.6 class) bill $10–60 per 1M output tokens, off this chart’s right edge by an order of magnitude.
First-party evaluation · retrieval & answer quality
RAG evaluation: 506 cases, end to end
Measured 2026-08-16 on dev. Corpus: 20 documents, 919 chunks, 9 clients. Test set: 506 cases, each with a verbatim evidence span from its source chunk.
| Metric | Result |
|---|---|
| 100.0% | |
| |
| IndicMSMARCO R@10 | 93.2% |
| |
| −4.8pp | |
| |
| 92.9% | |
| |
| unsupported claims | 7.9% |
| |
| defects | 1 |
| |
| −1.4pp | |
|