Gujarati, in beta
Available now. Send Gujarati text and you get Gujarati audio from the same speech endpoint, with the same voices and at the same price. You don’t add a language parameter or any other field, and you don’t change any code, because the model reads the script you send.- Send Gujarati in Gujarati script (ગુજરાતી), since that is what the model reads.
- English words stay in Latin script, mixed in mid-sentence, exactly the way they already work in Hindi.
- Quality is still being tuned. Beta means it is good enough to build against and not yet at the bar our Hindi voices hold. Listen before you put it in front of customers, and tell us what you hear. We tune the voices on beta feedback.
Why these six
Together these languages are the mother tongue of 429 million people, roughly every third Indian. For most of them, Hindi is no substitute. The census records how many speakers of each language also speak Hindi, and for four of the six the answer is almost none.
357 million people cannot be called in Hindi at all, so a voice agent that
speaks only Hindi cannot reach them. Only 1.5% of Tamil speakers report
speaking Hindi; for Telugu it is 5.7%, for Kannada 4.7% and for Bengali 8.6%.
With Hindi and these six together, an agent can speak the mother tongue of
about four in five Indians.
Speaker counts are first-language speakers from the 2011 Census of India;
“also speak Hindi” is Hindi reported as a second language (table C-17). These
are the most recent official figures. The next census is due in 2027.
What support means
- Same endpoint, same request shape. The speech endpoint does not change. You send text in the language’s own script, code-switched with English the way real conversation is, and get audio back.
- Native voices. Each language ships with its own voices in the voice gallery, evaluated by native speakers. We do not use an accented Hindi voice to read Tamil.
- TTS ships first. This page tracks the voice. A full agent conversation in a language also needs our speech recognition and LLM to clear the same bar, and that work follows the voice, so each language lands on the speech endpoint first.