IISc releases voice AI model for 65 Indian languages


IISc releases voice AI model for 65 Indian languages

Bengaluru: Researchers at the Indian Institute of Science (IISc) have released an open-source voice AI model that can recognise speech in 65 Indian languages and dialects.The model, called SraVaani, was developed by IISc’s SPIRE Lab with ARTPARK and support from Google. It can convert spoken words into text across 10 scripts and automatically identify the language being spoken, without users having to select one first.SraVaani supports 20 scheduled Indian languages and 45 regional languages and dialects, including Garo, Angika, Chakma, Kokborok, Tulu, Bundeli and Bajjika.This matters because many voice assistants and speech-to-text tools work well only in widely spoken languages. IISc said several of the languages covered by SraVaani are not officially supported by existing speech-recognition systems. Together, these languages are spoken by about 25 crore people, based on the 2011 Census.The model has been made freely available on Hugging Face under an MIT licence, allowing startups, researchers and developers to use, modify and build products on top of it.IISc said SraVaani performed on par with leading Indian speech-recognition systems on commonly supported languages. Its bigger advantage was in underserved languages. On Garo, for instance, it recorded a word error rate of 9.5%, compared with 69.4% for the next-best system tested. A lower error rate means the software makes fewer mistakes while converting speech into text.“When we began Project Vaani four years ago, the aim was simple: that voice AI should work for every Indian, not only for those whose languages already had the resources behind them,” said Prasanta Kumar Ghosh, professor at IISc and principal investigator of Project Vaani.SraVaani was trained using data collected under Project Vaani. The project has recorded more than 31,000 hours of speech from 156,000 people across 165 districts in 28 states. Participants spoke naturally instead of reading prepared sentences, helping capture accents, dialects and everyday speech more accurately.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *