Science & Technology

IIIT Hyderabad Study Reveals Why AI Can Misunderstand Spoken Telugu Questions

A new IIIT-H study using the VākQA benchmark shows how speech-recognition and translation errors can alter Telugu questions and reduce AI performance in spoken question answering.

IIIT Hyderabad Study Reveals Why AI Can Misunderstand Spoken Telugu Questions

IIIT Hyderabad Researchers Find Spoken Telugu Can Confuse AI, Causing Errors in Answers

Hyderabad: An AI system might give the wrong answer when asked the same question in speech compared to text. So concludes a new IIIT Hyderabad-led study, now underway to better understand the challenges of developing AI systems for Indian languages.

The Study finds that the quality of answers from a large language model can vary substantially depending on the input modality - in this case, whether the question is given in text or speech. The research introduces VākQA - a novel benchmark for spoken question answering in Telugu, comprising of 2,001 factoid question-answer pairs from six domains (science, general knowledge, politics, history, culture and geography) and 2.53 hours of spoken language, bilingual transcripts and human-verified reference answers. The paper was first shared on arXiv on September 17, 2026, with a public release in Hyderabad on September 21.

What Did the IIIT-Hyderabad Study Find?

The researchers set out to explore potential differences in the performance of AI-powered question-answering systems when presented with questions in different input modalities, including not only written language but also speech. The researchers found that on average, when presented with Telugu text, Gemini-2.5-Flash scored 3.63/5, while with Telugu speech, the score dropped to 3.28/5. In other words, the performance degraded by 21.2% on average across all questions when going from Telugu text to Telugu speech.

A similar phenomenon was observed when translating Telugu questions to English: the average score for Gemini dropped to 3.52/5 (down 18.8% compared to Telugu text). The researchers concluded that the underlying issue arises not due to a lack of information, but rather because the information can be changed or skewed before the model has a chance to act on it.

Speech Recognition Errors Can Lead To Dramatic Errors In Answer

The paper gives several compelling examples of how the error from speech recognition can cause the model to give vastly different answers than expected. In one instance, a question in Telugu about "our country's first satellite" was answered with Aryabhata (correct) when processed in Telugu, but after translation, the reference became ambiguous and the model answered with Sputnik.

In another example, a question about Telangana's state fruit was answered correctly as mango when processed in Telugu, but a speech recognition error that confused a word for "fruit" with a word for "festival" caused the question to be misunderstood and the answer changed to Bathukamma. A similar error occurred with a question about ringworm: the correct answer based on the Telugu text was fungus, but a speech recognition error in the question caused the model to give a different answer associated with heat.

The Error Chain Begins Before The Question Even Gets To The Model

The paper also notes that the researchers found considerable speech recognition errors (word error rate 30.25% and 35.23% respectively in the systems tested). When the combination of speech recognition, translation and question-answering is taken into account, there is potential for errors to accumulate at each stage of the processing pipeline. In the most error-prone speech-to-translation scenario reported in the study, Gemini's score dropped to 2.44/5 compared to 3.63/5 for Telugu text. This means that the voice assistant could be misled by the speech recognition or translation and give an answer that is only loosely related to the question the user intended to ask.

Why This Matters For Indian Languages And AI

The researchers note that evaluation of models on Indian languages often assumes that strong performance on English will translate to strong performance on Indian languages and spoken language tasks, but that this is not always the case. Telugu has cultural references and linguistic features that get lost or muddied in the translation or speech recognition process. VākQA aims to make these issues more visible rather than hiding them, and has been publicly released as part of this research, with the paper being accepted for presentation at IEEE Spoken Language Technology 2026.

For developers and researchers working on voice assistants, chatbots and other AI applications, the study offers practical insights about the importance of testing a model in the specific modality and language that the end user will be using, rather than simply relying on text or English as a proxy.

Final Thoughts

The IIIT-Hyderabad research does not suggest that AI systems are incapable of handling questions in Telugu, but rather that performance can degrade dramatically when a question is converted from text to speech, or from Telugu to English. VākQA provides a valuable benchmark for researchers to explore this phenomenon further. As voice enabled AI systems become more prevalent across India and its many languages, understanding the potential for speech recognition errors to compound with errors in the underlying language model may be of critical importance.

Click Here for More Science & Technology