POST /v1/audio/speech protocol, so every OpenAI SDK and every
framework with an OpenAI TTS integration can already call it. You change the
base URL and the model name, and the rest of your code stays as it is.
1. Get a key
Sign in to the console with your access code, open Developers, and create an API key. It is ask_live_ secret,
and the console shows it only once. We store only its hash, so we cannot
recover a lost key; you rotate it instead. The same key works for TTS and the
rest of the v2 API, and draws on the same wallet.
2. First audio
--output: without it, curl
prints 300 KB of binary into your terminal and you have to reset it. If
speech.wav plays, the integration works, and the sections below refine it.
There are exactly two formats:
Any other value, including OpenAI’s default
mp3, returns a 400 that names
these two. Always send response_format explicitly.
One model name.
model is mira-tts. The old build-numbered name,
mira-tts-v51, is still accepted as an alias, so existing callers keep
working. Any other value returns a 404.POST /v2/tts if you would rather keep
every URL under /v2. It runs the same handler and also accepts text as an
alias of input. Auth, fields, headers and billing are identical, so switching
between the two changes only the URL.
3. Pick a voice
ashu, neha, shruti and sameer. All four are Hindi
voices that code-switch into English the way real Indian speech does. neha is
the default if you omit voice. Each entry gives the voice_id, the voice’s
name and its language. To hear a voice before you choose, use the voice picker
in the console’s Agent Studio. An unknown voice returns a 400 that names the
allowed set; we never swap in another voice silently. Gujarati is accepted in
beta, and five more languages are on the way. See the
language roadmap.
4. Stream it with the OpenAI SDK
5. Drop it into your orchestrator
If you build on Pipecat, use our integration. Install it withuv add pipecat-mirai and add the service to your pipeline:
https://sandbox.voice.miraiminds.co/v1 with model mira-tts
and response_format="pcm", and treat the audio as 48 kHz.
What comes back
Every200 carries its own accounting in headers, so you can meter spend and
pin a build without a second request:
Every error after authentication also carries
X-Request-Id, so you can trace
a failure the same way as a success. Quote the id when you ask us about a
failed request.
Limits
We refuse a request over the limit instead of queueing it. Honour
Retry-After and send it again. A refused request does not hold a slot and
costs nothing. at_capacity can also come from a deployment-wide ceiling when
the speech node is saturated. Handle it the same way.
The limit is set per workspace, and a workspace starts at 10. The plans under
Billing list the concurrency each plan includes. If your
workspace’s limit doesn’t match your plan, or you need more, ask your Mirai
contact. Changing the limit is a provisioning change and needs no change to
your code.
Billing
₹1.80 per 1,000 characters. Top-ups use the same rate.
Shared endpoint. No contractual SLA. Monthly commitments are credited to your wallet. Unused credit never expires. All prices exclude GST.
Enterprise: Custom volume pricing, concurrency, dedicated endpoint and SLA. Top up anytime at the same rate.
Each request is rounded up to the paisa, and we charge only for completed
audio. If the engine fails, the stream stalls or your client disconnects
part-way, nothing is charged and no ledger row is written. The request still
appears in your log, with a status that says why.
The charge is one debit row on the same wallet your calls draw
from, and it appears on the console billing statement under Text to speech:
GET /v2/wallet/transactions
If your wallet cannot cover the request, you get a
402 before anything is
synthesised. See Wallet.
Your request log
We log every request, whether it completed, failed or was rejected, and you can read the log yourself.200 OK
GET /v2/tts/requests/{id} returns a single row. The id is the
X-Request-Id you got back from the request.
We do not keep the text you send.
preview is the first 48 characters of
input, so a row is recognisable in a list. The rest is never stored.When something fails
Errors use the same envelope as the rest of the API. See
Errors. If you are stuck on something this page does not cover,
reply on your onboarding thread, where a person reads it.