Saaras V4: Sarvam AI's 22-Language Speech Recognition Model

Sarvam AI is taking a more India-specific approach to speech recognition with Saaras V4, its latest automatic speech recognition model designed to work across 22 Indian languages and English. Sarvam announced Saaras V4 on August 24, and its documentation shows that the model became generally available on September 2.

It is built around a combination of an audio encoder and a 3-billion-parameter hybrid state-space language model that Sarvam says it trained completely from scratch in-house. The company’s pitch is not simply that the model recognises more languages.

Instead, Saaras V4 is designed around some of the problems that make Indian speech particularly difficult for automated systems: code-switching, dialect variation, noisy audio and multiple ways of representing the same spoken content. Since general availability, Sarvam has added Saaras V4 to its Realtime API and expanded keyterm prompting support across its speech endpoints, making the model part of an evolving production API.

Saaras V4 Supports 22 Indian Languages and Five Output Modes

The most immediately useful part of Saaras V4 is that it does not treat speech recognition as simply converting audio into words in the same language.

Representational image based on an official image | News

Sarvam gives developers five output modes: transcribe, translate, verbatim, translit and codemix. These modes can be applied to the same underlying audio input, although their supported output behaviour differs.

  1. Transcribe produces text in the original language and is the conventional speech-to-text option.
  2. Translate translates supported Indic-language speech into English. It is useful when the original conversation needs to be made accessible to an English-speaking user or downstream system, but it should not be understood as unrestricted translation between any two of the supported languages.
  3. Verbatim is intended to preserve how something was actually spoken, including features such as filler words and repetitions.
  4. Translit converts recognised speech into Roman script. This can be useful when a user speaks an Indian language but prefers reading it using the Latin alphabet.
  5. Codemix is designed for conversations in which languages are mixed naturally, such as Hindi-English speech. Instead of forcing every word into a single language, the output can retain the mixed-language character of the conversation.

Sarvam says keeping these capabilities inside one model also avoids additional processing stages that could introduce cascading errors.

Why the Architecture Is Different

The encoder processes the incoming speech signal, while the decoder generates the requested textual representation. Sarvam says the decoder is a 3-billion-parameter hybrid state-space language model trained from scratch by the company. That distinction matters.

Saaras V4’s 3B model should not be described as a standalone general-purpose LLM. It is the language-model component within a speech-recognition system, where its job includes interpreting acoustic information and generating the appropriate output format.

Sarvam’s decision to train the decoder from scratch also reflects the specific requirements of speech recognition. A decoder designed around multilingual speech has to deal with linguistic patterns that may not be well represented by a general text model, particularly when the target includes low-resource languages, code-mixed speech, and spoken-language irregularities. The result is an architecture designed around the speech task.

Designed for Noisy and Code-Mixed Indian Speech

Real-world Indian speech is rarely as clean as a laboratory recording. People speak in crowded environments, switch languages within the same sentence, use regional pronunciations, and move between formal and conversational vocabulary. Dialect variation can also make a single language considerably less uniform than a benchmark dataset might suggest.

Sarvam says it specifically trained and evaluated Saaras V4 against these conditions. The company describes the model as designed to improve robustness to noisy and code-mixed audio, while also addressing dialect variation and language identification. This matters for applications such as call centres, voice agents and field recordings, where users are unlikely to speak in carefully recorded, standardised sentences.

Sarvam reports a 2.9% language-identification error rate across the 10 most widely spoken Indian languages, compared with 5.22% across all 22 Indian languages, using verified IndicVoices data. These are Sarvam-reported language-identification error rates, not Word Error Rate figures, so they should not be interpreted as direct measurements of transcription accuracy.

Sarvam’s Benchmark Claims Show the Model’s Ambition

Sarvam reports that Saaras V4 achieved state-of-the-art performance across all 22 Indian languages in its evaluations.

For Indian-language testing, Sarvam evaluated the model using the Vistaar benchmark across 10 languages, using both conventional Word Error Rate and LLM-WER. WER measures transcription errors, while Sarvam describes LLM-WER as an additional semantic assessment intended to account for whether differences in transcription actually change the meaning of the content. Sarvam also evaluated Saaras V4 on seven English speech-recognition datasets covering areas including meetings, podcasts, financial calls, general speech, and different English accents.

Saaras V4
Representational image based on an official image | News

The company reports that Saaras V4 recorded the lowest average Word Error Rate across those seven datasets, six of which represent foreign English dialects and one of which represents Indian English. These are Sarvam’s reported benchmark results. The datasets, evaluation setup, and comparison models matter when interpreting any claimed state-of-the-art or lowest-WER result.

Real-Time Speech Recognition Is Part of the Product

Saaras V4 is also intended for applications where waiting for an entire recording to finish is not practical. Sarvam reports a time to first token below 150 milliseconds for streaming applications. The model can provide partial results through its WebSocket-based real-time infrastructure, making it suitable for voice agents and live transcription.

Its current API availability is broader than a simple streaming endpoint. Sarvam provides REST, Batch, and WebSocket options. REST supports synchronous transcription for shorter recordings, while the Batch API can process files up to two hours and supports speaker diarization. WebSocket options provide streaming transcription for live applications. Developers can access Saaras V4 through Sarvam’s Python and Node.js SDKs, with integrations for platforms and frameworks including Vercel AI SDK, LiveKit Agents and Pipecat Agents. The September updates also show that the model is being developed as a production API.

Sarvam added Saaras V4 to the Realtime API on September 3. Keyterm prompting was subsequently introduced for REST, Batch, and WebSocket use, and the September 22 update extended keyterm prompting to both the legacy WebSocket and Realtime WebSocket endpoints. Keyterms allow developers to bias recognition toward names, places, brands, or technical terms without guaranteeing that those terms will appear in the output.

However, Saaras V4 is not automatically the default model across Sarvam’s speech APIs. Saaras V3 remains the default or recommended model in Sarvam’s current documentation, including for the Realtime API. Developers must explicitly select Saaras V4, where supported, using saaras:v4.

Saaras V4 vs Saaras V3

Saaras V4 builds on the same basic purpose as V3. Both versions support the five output modes and the same 22 Indian-language coverage. The most notable expansion reported for V4 is English: while V3 focuses on Indian English, V4 extends recognition to Global English accents as well.

Saaras V4
Representational image based on an official image | News

There is also an important product-status detail. Saaras V3 remains the default recommended model in Sarvam’s current documentation, while V4 is the newer model and can be explicitly selected saaras:v4. The Realtime API also lists V3 Realtime as its default while making V4 available as an option. That means “latest model” and “default model” are not currently synonymous within Sarvam’s API stack.

What Saaras V4 Could Mean for Indian Voice AI

The practical significance of Saaras V4 is ultimately less about the number 22 and more about how many speech problems it attempts to address within one system. A multilingual call-analysis platform could transcribe a Bengali conversation, identify code-mixed English terms, and preserve speaker separation. A voice agent could process Hindi-English speech in real time. A regional-language application could offer native-script transcription while providing Romanised output for users who are more comfortable reading Latin script.

Those workflows are particularly relevant in India because language choice is often fluid. A speaker can change languages, mix English terminology into an Indian-language sentence, use regional vocabulary, and speak through a low-quality phone connection, all within the same interaction. That is the problem Saaras V4 is attempting to solve.

Sarvam’s own benchmark claims still need to be understood in their evaluation context, but the architecture and API design point to a broader goal: making speech recognition useful across the messy conditions of real multilingual communication. With 22 Indian languages, Global English support, five output modes, streaming recognition and production APIs, Saaras V4 is positioned less as another transcription model and more as infrastructure for building voice applications around India’s linguistic diversity.

(Source)

Leave a Comment