How well do speech APIs actually hear India?

Seven speech-to-text engines, tested on the same audio: read speech in 12 Indian languages, Hinglish conversations, noisy phone calls, and fresh real-world recordings no model has trained on. Every test case is inspectable below, ground truth against output, word by word.

Leaderboard

Accuracy = words transcribed correctly, after normalizing away script choice (writing "start" vs "स्टार्ट" for the same spoken word is not an error). Computed on the exact same clips for every engine.

Accuracy by language

Average per-clip accuracy for each engine on each test type. Darker teal = better. This is the map of who to route where.

The shape of the gap

Accuracy by test type

read speech vs Hinglish vs phone calls vs fresh real-world audio

Accuracy vs speed

up and left is better · bubble size = cost per audio hour

What an hour of audio costs

measured from per-call billing or list price

Processing speed

seconds of processing per second of audio (lower = faster)

Every test case, inspectable

Gramvaani phone-call cases are scored in the leaderboard but not shown here: that corpus is licensed for academic use only and does not permit redistributing its transcripts.

Ground truth on top, every engine's output below it. Differences are highlighted: wrong word added word missed word. Click any case to expand.

Method, honestly

Two kinds of tests

Standard datasets (FLEURS read speech, MUCS Hinglish, Gramvaani phone calls) are public; providers have almost certainly trained on them, so treat those scores as a flattering baseline. Real-world 2026 is Creative-Commons YouTube audio published January–June 2026, after every model's training cutoff. Nobody could have memorized it. When the two disagree, trust the real-world number.

Where ground truth comes from

Standard datasets ship human transcripts. Real-world references come from multi-model consensus with a strict agreement gate; the 36 of 132 chunks that failed the automatic gate were adjudicated case by case (accept only on cross-model content agreement, every decision and rationale logged in the repo). Sarvam saarika, Gemini, and Whisper contributed consensus hypotheses, so their real-world scores are upper bounds (marked); the per-system exposure is quantified as self_ref_rate in the data. Saaras v3, gpt-audio, and Voxtral never saw the references: their scores are fully independent.

The script problem

Indian speech mixes languages mid-sentence, and engines legitimately disagree on script: "फॉन्ट साइज" and "font size" are the same spoken words. Headline accuracy romanizes both sides before comparing so script choice is not punished. The diff view still shows the raw outputs, so you can see script differences without being misled by them.

Models as run