Science & Technology

IIT Madras Bodhan AI Launches 4 Open AI Models for Indian Languages

IIT Madras-incubated Bodhan AI and AI4Bharat have launched four open AI models for Indian languages, covering speech recognition, text-to-speech, translation and OCR across up to 27 languages.

IIT Madras Bodhan AI Launches 4 Open AI Models for Indian Languages

IIT Madras Bodhan AI Launches 4 Open AI Models for Indian Languages

Chennai: Artificial Intelligence in India is breaking away from its English-only model. Bodhan AI, an AI in Education Centre of Excellence at IIT Madras, has collaborated with AI4Bharat to create four open-source AI models for Indian languages, prompting AI model adoption in education, research, government and tech sectors.

Speech-to-text conversion, text-to-speech conversion, machine translation, and Optical Character Recognition (OCR) models together are strong building blocks of AI models. However, their value is only noticed when they are integrated. The move is also part of the Bharat EduAI Stack, which Bodhan AI describes as sovereign digital public infrastructure for India's education system. The Centre of Excellence itself was launched by the Minister of Education, Dharmendra Pradhan, in February 2026.

Four AI Models, Four Core Capabilities

The new models are designed to handle different parts of the multilingual AI problem.

Model

Capability

Language Coverage

Indic-Transcribe

Automatic speech recognition

27 languages

Indic-Speak

Text-to-speech

23 languages

Indic-Translate

Machine translation

22 languages

Indic-OCR

Optical character recognition

23 languages

Indic-Transcribe converts spoken language into text, while Indic-Speak does the reverse by generating speech from written text. Indic-Translate enables translation between Indian languages, and Indic-OCR can identify and extract text from documents and images.

That combination could be particularly useful in classrooms. A student could speak a question in an Indian language, an education application could process textbook or worksheet content through OCR, and the resulting material could be translated or read aloud.

Why the Open-Model Approach Matters

Bodhan AI is not releasing the models as end-user AI services. Rather, they are releasing them as unified infrastructure for others to build further upon. These models provide AI infrastructure for speech, translation, and OCR to edtech companies, startups, researchers, universities, government organisations, and technology companies.

AI4Bharat has extensive AI model-building experience. IIT Madras's research lab has developed open-source Indian language datasets for translation, speech, and speech-to-text.

NVIDIA Technologies Used in the Models

To build and optimise models, Bodhan AI trained and post-trained the NVIDIA Nemotron open models and the NVIDIA NeMo library. To further improve Indian language models, they also supported post-training of NVIDIA Nemotron 3.5 ASR. Inference for the models is done using the NVIDIA TensorRT-LLM and vLLM microservices.

The collaboration includes datasets, training recipes, and evaluation methods for future Indian language foundational models.

AI Tutor for Classes 6 to 12 Also Launched

Bodhan AI has developed beyond just models. With Student Tutor Bot, Bodhan AI targets students from Classes 6 to 12 and is built on the NCERT and SCERT curriculum. Students can interact with the tutor in 22 languages using either text or voice. The system has the capability to provide students with instructions and examples, perform evaluations, and interact with the students to aid them in solving problems.

A Teacher Assistant Bot has been developed as well. This new Bot will assist teachers in the creation of lesson plans, worksheets, quizzes, homework and revision materials, based on given parameters, such as grade, subject, topic, duration, and difficulty level. The assistant is built on top of AI; however, the teacher has control over all AI-generated materials.

A Push Towards Sovereign AI In Education

Bodhan AI seeks to develop an infrastructure for education that is flexible enough for multiple languages, generalizable and sovereign. Its own website credits the infrastructure as a base layer for the AI system for education in India that supports learning, tutorials, practice, and assessments.

The launch of the four models in September 2026 is not just about the release of four models. With developers potentially adopting them, the technology has the potential to lessen duplication in India's education system and make it easier to construct AI tools surrounding the country's diverse languages. Currently, the true challenge is not if India can improve the release of Indian language models. The real challenge is if developers, schools, and public institutions can develop those models into tools that students can use in their own languages.

Click Here for More Science & Technology