News

Researchers have released SraVaani-1.0, described as the first openly evaluated multilingual Indian automatic speech recognition (ASR) system trained on 65 Indian languages and dialects. Many of the languages and dialects it covers, including several tribal and low-resource ones, currently have no other publicly available speech-recognition system.

About SraVaani-1.0

  • What It Is: SraVaani-1.0 is an automatic speech recognition (ASR) model that converts spoken Indian languages and dialects into text, built on the FastConformer architecture.

  • Training Data: It was pre-trained on 31,255 hours of unlabelled speech from the VAANI corpus, then fine-tuned on 31,263 hours of labelled speech drawn from 24 public datasets.

  • Working: It uses a three-stage training pipeline with a Hybrid Token-and-Duration Transducer (TDT)-CTC decoder, including a step that aligns speech with paired images to build richer, meaning-based representations for low-resource languages.

  • Significance: It is described as the only open-source, openly evaluated ASR model that can transcribe several low-resource and tribal Indian languages, matching or beating existing multilingual systems on standard benchmarks.