Falcon Benchmark

Frontier-class Indic AI at 1/13th the serving cost.

A 26B-parameter sparse mixture-of-experts, about 4B active per query, FP4-quantized and benchmarked exactly as it shipped on a single GPU. It outscores every same-class open model measured here across 11 Indian languages, tracks models many times its size on English reasoning, and serves at $0.42 per million tokens.

  • One pinned pipeline
  • Greedy decoding
  • Fixed seeds
  • Full per-sample audit trail
  • Every number re-runnable
$0.42

per 1M output tokens

13× cheaper than Llama-70B, 5× than Gemma-3-27B, measured on the same GPU

82.0

MMLU-Pro

within 3 pts of DeepSeek-V3.1’s published 84.8, at about 4B active params

83.2

MILU Indic composite (gen-CoT)

best in panel across 10 Indian languages

2,029

tokens/sec @ 30 ms TTFT

12 concurrent streams, production serving stack

Serving economics: measured, identical hardware

BenchmarkFalconQwen3-30B-A3BGemma-3-27BMistral-Small-3.2Sarvam-MLlama-4-Scout 109BLlama-3.3-70B
TTFT ms @ 12512-tok prompt · lower better30.0038.0048.0040.0050.0043.0059.00
Decode tok/s12 concurrent2,0291,699388480479921152
$ per 1M output tokenslower better0.420.502.171.751.760.915.54

Measured on one RTX PRO 6000 (Modal $3.03/h), identical load.

Indic languages: composite scores

BenchmarkFalconQwen3-30B-A3B
MILU (generative CoT)avg 10 langs, India-centric MCQ83.273.8
MMLU-ProX hard MCQavg hi/bn/mr/te, CoT72.955.4
GSM8K-Indic mathavg 10 langs, CoT88.385.4
GSM8K romanized HindiHinglish stressor86.077.2
Translation en→Indicavg 10, chrF++50.141.2
Translation Indic→enavg 10, chrF++61.658.0
Hinglish→EnglishCOMI-LINGUA, 3 refs, chrF++79.376.0

Averages over the product’s Indian languages (10 scripts + romanized Hindi). MCQ composites in %, translation in chrF++. Per-language breakdowns: results_flat.csv.

Global languages: Falcon composite scores (13 locales)

Falcon, measured
BenchmarkFalcon
Belebele reading comp.avg 13 global langs56.1
MMLU-ProX hard MCQavg 9 langs, CoT75.7

zh, ja, ko, es, de, ru, sv, pl, it, nl, ro, ar, id. Falcon’s own measured scores across the 13 non-Indic locales the product supports; no competitor is run on this suite, so no row is marked best.

English core

BenchmarkFalconQwen3-30B-A3BMistral-Small-3.2
MMLU-Pro5-shot CoT82.077.963.2
GSM8K8-shot CoT94.092.183.9
MATH-500math-verify95.296.287.6
AIME 2025flex76.770.036.7
IFEvalprompt-strict89.882.674.5
HumanEvalpass@199.490.988.4
MBPP+pass@196.692.984.9

All cells measured by this pipeline.

Scale contrast: bigger Meta models on the same GPU

Serving rows measured on identical hardware
BenchmarkLlama-4-Scout 109BLlama-3.3-70BFalcon
Total params (B)109.070.626.0
Active params (B)17.070.64.0
Decode tok/smeasured920.8151.92,028.9
$ / 1M tokensmeasured0.95.50.4

Serving rows measured by this pipeline on the same GPU. This table makes a cost argument only: 4.2× the parameters at 2.2× the cost (Scout) or 13× the cost (70B). Quality is deliberately not compared here: the full quality suites for these two models were not run, and setting our measured scores against vendor-published numbers taken under different protocols would not be a like-for-like comparison.

Quality per dollar

MMLU-Pro vs $/1M tokens · log scale
MMLU-Pro accuracy versus measured serving costFalcon scores 82.0 at about 42 cents per million output tokens, the highest measured accuracy at the lowest measured cost among the models plotted here.$0.5$1$2$5406080$ per 1M output tokens (log scale) →MMLU-Pro →Qwen3-30B-A3BMistral-Small-3.2Llama-4-Scout 109B°Llama-3.3-70B°Falcon

MMLU-Pro accuracy vs measured serving cost (log scale). Solid = both axes measured by this pipeline; ° = vendor-reported quality with our measured cost. Falcon sits at the top-left of the plotted set: the highest measured accuracy at the lowest measured cost among these models. Frontier APIs (GPT-5.6 class) bill $10–60 per 1M output tokens, off this chart’s right edge by an order of magnitude.

First-party evaluation · retrieval & answer quality

RAG evaluation: 506 cases, end to end

measured 2026-08-16 · dev

Measured 2026-08-16 on dev. Corpus: 20 documents, 919 chunks, 9 clients. Test set: 506 cases, each with a verbatim evidence span from its source chunk.

MetricResult
100.0%
IndicMSMARCO R@1093.2%
−4.8pp
92.9%
unsupported claims7.9%
defects1
−1.4pp