How well do speech APIs actually hear India?
Seven speech-to-text engines, tested on the same audio: read speech in 12 Indian languages, Hinglish conversations, noisy phone calls, and fresh real-world recordings no model has trained on. Every test case is inspectable below, ground truth against output, word by word.
Leaderboard
Accuracy = words transcribed correctly, after normalizing away script choice (writing "start" vs "स्टार्ट" for the same spoken word is not an error). Computed on the exact same clips for every engine.
Accuracy by language
Average per-clip accuracy for each engine on each test type. Darker teal = better. This is the map of who to route where.
The shape of the gap
Accuracy by test type
Accuracy vs speed
What an hour of audio costs
Processing speed
Every test case, inspectable
Gramvaani phone-call cases are scored in the leaderboard but not shown here: that corpus is licensed for academic use only and does not permit redistributing its transcripts.
Ground truth on top, every engine's output below it. Differences are highlighted: wrong word added word missed word. Click any case to expand.