Hindi as it is spoken
Code-mixed Hindi and English, with names and numbers mid-sentence. Tuned on conversational, telephone-band Hindi.
Mira STT · Speech to text
Mira STT turns Hindi, and the English mixed into it, into text. On RinggAI’s public benchmark of 10,000 Hindi utterances it made fewer word errors than Ringg, ElevenLabs, Sarvam and Deepgram.
Word error rate
RinggAI Hindi benchmark · 10,000 utterances
Lower is better
Code-mixed Hindi and English, with names and numbers mid-sentence. Tuned on conversational, telephone-band Hindi.
On Lahaja, 132 speakers from 83 districts it never trained on: 11.63% WER, against 16.55% for the open Qwen3-ASR model.
Our finetune of SraVaani-1.0, the open model from ARTPARK at IISc, served on our own inference engine.
Benchmark
| Dataset | Mira STT | Ringg Parrot | ElevenLabs | Sarvam | Deepgram |
|---|---|---|---|---|---|
| CommonVoice1,727 utterances | 12.78 | 12.62 | 15.98 | 16.68 | 18.49 |
| FLEURS418 utterances | 9.89 | 10.36 | 9.74 | 13.40 | 15.65 |
| IndicTTS100 utterances | 5.31 | 5.21 | 10.57 | 10.32 | 8.65 |
| Kathbath1,929 utterances | 7.99 | 9.46 | 12.14 | 13.32 | 14.18 |
| Kathbath noisy1,929 utterances | 9.47 | 10.84 | 11.83 | 14.49 | 15.61 |
| MUCS3,897 utterances | 8.24 | 8.41 | 8.97 | 10.50 | 16.03 |
| Overall10,000 utterances | 9.12 | 9.72 | 11.11 | 12.83 | 15.79 |
Normalised folds spellings that sound the same (अन्दर = अंदर, लिये = लिए), never grammar: है and हैं still count as different words. Vendor scores use the outputs bundled in the dataset; their model versions are Ringg’s.
The dataset on Hugging FaceOur Hindi speech-to-text model: a finetune of SraVaani-1.0, the open model from ARTPARK at IISc, tuned for conversational, telephone-band Hindi and served on our own engine.
On all 10,000 utterances of RinggAI’s public Hindi benchmark, Mira STT scored 9.12% word error rate, the lowest of the five systems scored. Ringg Parrot scored 9.72%.
Hindi, including English words mixed into Hindi speech. More Indian languages follow as each clears our own evaluation.
It is tuned on telephone-band audio and runs inside our voice agents today. The benchmark on this page is public read and prompted speech, not phone calls, so we do not quote a phone-call accuracy figure from it.
Try it inside a voice agent in the sandbox today. The standalone speech-to-text API is in early access: book a call and tell us about your audio.
Tell us what your callers sound like. We will run Mira STT on a sample of your audio before you commit.