Skip to main content
The streaming endpoint carries a whole conversation on one connection. You send text as your language model writes it, a token or a sentence at a time. We cut it into sentences, synthesise them in order and send the audio back on the same socket. Voices, output formats and the price are the same as on POST /v1/audio/speech.

When to use it

Use the WebSocket for live conversations: voice agents on the phone or in a browser, anything where a model writes the reply while someone waits for it. Compared with one HTTP request per sentence, it changes four things:
  • You pay for the connection once per call. A new HTTPS connection needs a TCP and TLS handshake, which takes 0.4 to 1 second from India and sometimes longer when a connection attempt has to be retried.
  • You can send tokens as they arrive. We find the sentence ends.
  • The next sentence is synthesised while the current one plays, so there is no pause between sentences.
  • Barge-in is one message. The sentence that was playing stops at once and is not charged.
When your workspace is at its concurrency limit, a sentence on a socket waits up to 3 seconds for a free slot. An HTTP request gets a 429 straight away. For one-off audio, such as a file, a voicemail or a notification, the HTTP endpoint is simpler.
If your voice agent runs on Pipecat, MiraiWebsocketTTSService from our Pipecat package speaks this protocol for you, including barge-in, capacity retries and reconnects.

Connect

wss://sandbox.voice.miraiminds.co/v2/tts/stream is the same endpoint, if you keep every URL under /v2. Open one socket per call and keep it until the call ends.

From your server

Send your API key in the Authorization header, as on every other request:
We refuse an API key in the URL (?api_key= returns 401), because URLs end up in proxy and server logs.

From a browser

A browser cannot set headers on a WebSocket, and your API key must never reach a web page. Your backend mints a short-lived stream token with its key and hands the page only the token.
201 Created
url is a path. The page connects to your API host followed by that path, as in the browser example. A token:
  • lasts 60 seconds and opens one socket;
  • bills the workspace of the key that minted it, under that key’s rate limit;
  • stops working if that key is revoked.
A socket opened with a token can synthesise speech on your wallet until it closes, for up to an hour. Mint tokens only for users you have signed in, one per connection, right before the page connects, and rate-limit the route that mints them.

When the connection is refused

We check everything we can before the WebSocket upgrade. A refused connection gets an ordinary HTTP response with the usual error envelope: A successful upgrade carries X-Session-Id: ttsws_…. Quote it when you ask us about a socket.

How a call flows

  1. Connect. The first message you receive is session.ready, with the settings in force and your limits.
  2. Send session.update to pick a voice and an output format. You can skip it if the defaults suit you.
  3. For each turn of the conversation, pick a new context_id and send the reply’s text in text messages as your model produces it.
  4. When the reply is complete, send flush.
  5. Each sentence arrives as audio.start, then its audio, then audio.done. After the last sentence of the turn you get context.done.
  6. If the caller interrupts, send cancel for that context. See Barge-in.
  7. Keep the socket for the next turn, and close it when the call ends.

Messages you send

Every message is a JSON text frame with a type. A message we cannot accept is answered with an error event and changes nothing. The socket stays open.

session.update

We check every field before applying any of them, so a bad value leaves all your settings as they were. If you change response_format without sample_rate, you get that format’s default rate. If response_format includes a rate, leave sample_rate out or send the same value. Audio for sentences already cut keeps the settings it was cut with, and every audio.start tells you what its audio is. wav is refused because a WAV header has to state the length of audio that does not exist yet. Ask for pcm at the rate you need.

text

A context is one reply, usually one turn of the conversation. context_id is any string of 1 to 128 characters that you choose. Send the text exactly as your model streams it, spaces included: we join the pieces, so a token without its leading space runs into the word before it.

Events you receive

session.ready
audio.done

Audio

Binary frames carry raw audio with no header, always in whole samples:
  • pcm is signed 16-bit little-endian mono (encoding: pcm_s16le), 2 bytes per sample.
  • μ-law and A-law are G.711, 1 byte per sample. At 8 kHz that is 8,000 bytes a second, so 160 bytes is 20 ms.
A binary frame always belongs to the most recent audio.start. Use audio_transport: base64 only if your client cannot handle binary frames: the same audio then arrives in audio.chunk events and takes about a third more bytes.

How text becomes sentences

We hold each context’s text until one of these happens:
  • A sentence ends: ., ?, !, ।, ॥ or …, optionally followed by a closing quote or bracket, and then whitespace.
  • A new line arrives.
  • The text passes 250 characters (limits.max_segment_chars) with no sentence end. We cut at the last , ; : — –, or else at the last space.
  • You send flush.
The rule needs whitespace after the stop, so 3.5 and example.com are never split, and a stop at the very end of what you have sent waits for your next token or a flush. A full stop after an abbreviation such as Rs., Dr., Mr. or e.g., or after a single initial, does not end a sentence. A sentence shorter than two words joins the next one, so “Okay.” is not sent on its own. Lengths are counted in Unicode code points, which is also how characters are billed.

Ordering

These hold on every socket:
  • Sentences are spoken one at a time, in the order they were cut, across all contexts.
  • Each audio.start is closed by exactly one audio.done or error with the same request_id, or by context.cancelled for its context. The next audio.start comes after that.
  • So every binary frame belongs to the most recent audio.start.
  • A context’s context.done comes after its last sentence’s audio.done or error.
  • Nothing for a context is sent after its context.cancelled.

Barge-in

When the caller starts talking over your voice agent:
  1. Send cancel with the context_id that is speaking.
  2. Drop any audio you still receive for that context. Some may already be on the wire. Track the speaking context from the latest audio.start, and start playing again at the next audio.start for a different context.
  3. Clear the audio you have already handed to your player or phone provider. On Twilio, send a clear event; on Plivo, clearAudio; on Exotel, clear.
  4. Use a new context_id for the next reply.
On our side, cancel discards the context’s unsent text and its queued sentences, and stops the sentence being synthesised straight away. None of it is charged. You always get context.cancelled, even if the context had already finished, so you never need to know whether you were too late.

When your workspace is busy

Each sentence takes one synthesis slot from your workspace’s concurrency limit, the same pool HTTP requests use, while it is synthesised. An open socket that is not speaking holds no slot. When a sentence’s turn comes and the workspace is full, it waits for a slot in a queue shared by all your sockets, for up to 3 seconds. If no slot frees in that time, you get an error with code at_capacity and retry_after_secs for that sentence, and the socket moves on to the next one. Send that text again if you still need it. pipecat-mirai does this once for you. If you run many calls at once, ask your Mirai contact for a higher limit. Ask for a little more than your peak number of simultaneous calls, because each call starts its next sentence shortly before the current one finishes.

Billing

The socket costs the same as the HTTP endpoint, per character of text spoken:
  • A sentence is charged when its audio.done is sent, and cost_paise in that event is exactly what was debited.
  • A sentence that is cancelled, fails, or is still undelivered when the socket closes is not charged.
  • Fractions of a paisa carry over from one sentence to the next, so the total for a socket is its characters times the rate, rounded up once. A streamed reply never costs more than sending the same text in one HTTP request, and the spaces between sentences are not billed.
  • Because of that carry, some sentences cost 0 paise. They appear in your request log but get no wallet row.
Every sentence has its own row in your request log, with transport set to websocket. Its request_id is the row’s id and the reference on its wallet row. Before each sentence we check that your wallet can cover it. If it cannot, that sentence fails with insufficient_balance.

Limits

We send a WebSocket ping every 20 seconds and drop a client that has not answered for 60 seconds. Most WebSocket libraries, and every browser, answer pings for you. Pings do not count as activity for the idle limit: to keep a quiet socket open, send {"type": "session.update"} with no fields every 30 seconds or so. It changes nothing. If your client stops reading, we stop reading your messages until it catches up, and a client that stays stuck is disconnected. Read the socket all the time, and do slow work such as pacing audio to a phone line on another task.

Errors on an open socket

An error event affects one message or one sentence. The socket stays open and the next sentence goes ahead.

Closing

Close the socket yourself when the call ends. If you send close first, we finish speaking what you have sent before closing. When we close a socket, we send session.closed with a reason, then close it: A sentence that was not delivered when a socket closed is not charged.

Python example

This script streams a model’s reply into the socket token by token and saves 8 kHz μ-law, the format a phone line takes. Replace llm_tokens with your model’s stream.
stream_tts.py
With websockets older than 14, pass extra_headers= instead of additional_headers=.

Bridge to a phone call

To play the audio on a Twilio call, forward each binary frame to the Media Stream instead of writing it to a file. The class below does that for one call. It keeps Twilio at most 0.4 seconds ahead of what the caller hears, which absorbs short stalls on your server without making barge-in slow, and it handles barge-in. Plivo and Exotel work the same way with their own message names, and Exotel takes pcm_8000 instead of μ-law (see For phone calls).
phone_leg.py
Read the Mirai socket in its own task and pass each message to on_mirai_message. The pacing then waits in that task while the rest of your voice agent keeps running. The same pacing is what apply_output_lead does in Pipecat.

Browser example

A page that speaks typed text, with a Stop button for barge-in. Your backend mints the token, so the API key stays on your server. requireSignedIn stands in for your own sign-in check.
server.js
The page plays 24 kHz PCM through an AudioWorklet and holds 150 ms of audio before each sentence starts, so a short first chunk cannot start playback and then stall.
index.html
A token opens one socket, so connect mints a new one each time it needs to reconnect, for example after the socket has been idle for two minutes.

TTS quickstart

The HTTP endpoint, output formats, billing and the request log.

Pipecat

MiraiWebsocketTTSService and the recommended phone setup.

Fixing choppy audio

Symptoms, causes and fixes for gaps and breaks.

Limits

Concurrency, sockets and rate limits.