Skip to main content
This is the voice our agents use, offered as a standalone API. It speaks OpenAI’s POST /v1/audio/speech protocol, so every OpenAI SDK and every framework with an OpenAI TTS integration can already call it. You change the base URL and the model name, and the rest of your code stays as it is.

1. Get a key

Sign in to the console with your access code, open Developers, and create an API key. It is a sk_live_ secret, and the console shows it only once. We store only its hash, so we cannot recover a lost key; you rotate it instead. The same key works for TTS and the rest of the v2 API, and draws on the same wallet.

2. First audio

The response body is the audio itself. Keep --output: without it, curl prints 300 KB of binary into your terminal and you have to reset it. If speech.wav plays, the integration works, and the sections below refine it. There are exactly two formats: Any other value, including OpenAI’s default mp3, returns a 400 that names these two. Always send response_format explicitly.
One model name. model is mira-tts. The old build-numbered name, mira-tts-v51, is still accepted as an alias, so existing callers keep working. Any other value returns a 404.
The same endpoint is also mounted at POST /v2/tts if you would rather keep every URL under /v2. It runs the same handler and also accepts text as an alias of input. Auth, fields, headers and billing are identical, so switching between the two changes only the URL.

3. Pick a voice

The voices today are ashu, neha, shruti and sameer. All four are Hindi voices that code-switch into English the way real Indian speech does. neha is the default if you omit voice. Each entry gives the voice_id, the voice’s name and its language. To hear a voice before you choose, use the voice picker in the console’s Agent Studio. An unknown voice returns a 400 that names the allowed set; we never swap in another voice silently. Gujarati is accepted in beta, and five more languages are on the way. See the language roadmap.

4. Stream it with the OpenAI SDK

5. Drop it into your orchestrator

If you build on Pipecat, use our integration. Install it with uv add pipecat-mirai and add the service to your pipeline:
It handles sample rates and interruptions for you, and has a one-line fix for choppy audio on phone calls. See Pipecat for the full setup. Other frameworks with an OpenAI TTS integration can call the endpoint directly. Point them at https://sandbox.voice.miraiminds.co/v1 with model mira-tts and response_format="pcm", and treat the audio as 48 kHz.

What comes back

Every 200 carries its own accounting in headers, so you can meter spend and pin a build without a second request: Every error after authentication also carries X-Request-Id, so you can trace a failure the same way as a success. Quote the id when you ask us about a failed request.

Limits

We refuse a request over the limit instead of queueing it. Honour Retry-After and send it again. A refused request does not hold a slot and costs nothing. at_capacity can also come from a deployment-wide ceiling when the speech node is saturated. Handle it the same way. The limit is set per workspace, and a workspace starts at 10. The plans under Billing list the concurrency each plan includes. If your workspace’s limit doesn’t match your plan, or you need more, ask your Mirai contact. Changing the limit is a provisioning change and needs no change to your code.

Billing

₹1.80 per 1,000 characters. Top-ups use the same rate. Shared endpoint. No contractual SLA. Monthly commitments are credited to your wallet. Unused credit never expires. All prices exclude GST. Enterprise: Custom volume pricing, concurrency, dedicated endpoint and SLA. Top up anytime at the same rate. Each request is rounded up to the paisa, and we charge only for completed audio. If the engine fails, the stream stalls or your client disconnects part-way, nothing is charged and no ledger row is written. The request still appears in your log, with a status that says why. The charge is one debit row on the same wallet your calls draw from, and it appears on the console billing statement under Text to speech:
GET /v2/wallet/transactions
If your wallet cannot cover the request, you get a 402 before anything is synthesised. See Wallet.

Your request log

We log every request, whether it completed, failed or was rejected, and you can read the log yourself.
200 OK
Rows come newest first, and only from your own workspace: GET /v2/tts/requests/{id} returns a single row. The id is the X-Request-Id you got back from the request.
We do not keep the text you send. preview is the first 48 characters of input, so a row is recognisable in a list. The rest is never stored.
The same table is in the console under TTS → Requests, if you would rather check it there than poll the API.

When something fails

Errors use the same envelope as the rest of the API. See Errors. If you are stuck on something this page does not cover, reply on your onboarding thread, where a person reads it.