OpenAIRealtimeLLMService, or any other client for that protocol, you
change the API key and the URL and keep everything else:
Connect
prod.voice once
your company has gone live, sandbox.voice.miraiminds.co with a sandbox key.
A malformed parameter fails the upgrade with
400 invalid_request, and
error.param names the parameter.
Pipecat builds its URL as base_url + "?model=…". A base_url that already
carries a query, such as …/v2/realtime?agent_id=agt_…, therefore reaches us as
?agent_id=agt_…?model=…. That is fine: we read your parameters correctly.
Authentication
Send your workspace API key in theAuthorization header, as every OpenAI
Realtime client does:
openai-insecure-api-key.sk_live_…, alongside realtime. We answer with
the realtime subprotocol.
With Pipecat
Stock Pipecat 1.8.1 works unchanged. Here is a complete microphone-and-speaker agent. Save it aspipecat_realtime.py, put on headphones and run it with
uv, which installs Pipecat for you. The microphone
needs PortAudio (brew install portaudio on macOS).
pipecat_realtime.py
pipecat_realtime.py
- Send and play 16-bit PCM at 24 kHz. Set
audio_in_sample_rate=24000andaudio_out_sample_rate=24000inPipelineParams. - Remove your own STT, LLM and TTS services. The Realtime service replaces all three, and your transport and context aggregators stay as they are.
LLMRunFrame
(or send response.create) when you want it to speak first.
Sessions
- You open the socket. We check your key and balance and start a call on the
realtimechannel. - We send
session.created. - You send
session.updatewith your settings. We apply them and answersession.updated. We wait up to 1.5 seconds aftersession.createdfor this first update. If it does not arrive in time, the session starts with the agent’s settings (or the defaults), and your update is applied when it comes, apart from the language and audio formats. - You stream audio with
input_audio_buffer.append. We detect when the caller has finished, transcribe, run the model and stream the reply as audio. Audio you send before the session is ready is kept, up to 2 seconds of it. - You close the socket, or the session reaches its maximum duration. The call ends and is billed.
session.update again at any point. Most changes apply straight
away; session settings lists the exceptions.
Every event travels as a JSON text frame, audio included (base64). Binary
frames are ignored.
Events
The events are the OpenAI Realtime GA events. These are the ones we handle. You sendoutput_audio_buffer.clear and transcription_session.update are accepted and
ignored. Any other event type returns an error with code invalid_event.
We send
Not supported yet
rate_limits.updated.- Interim transcripts. Transcription deltas carry final segments only, so a client that joins the deltas never prints a word twice.
- Text-only replies. Every reply is spoken, so there is no
response.output_text.delta. - Per-response settings in
response.create. - Image and audio content in
conversation.item.create. Only text parts are read.
Session settings
“First update only” means the
session.update that arrives within 1.5 seconds
of session.created. A later change to the language or an audio format is
ignored, because both are fixed once the session has started.
Fields we do not know are ignored. Beta-shaped fields are accepted too:
input_audio_format and output_audio_format (pcm16, g711_ulaw,
g711_alaw), input_audio_transcription, turn_detection, voice, speed,
modalities, max_response_output_tokens and a top-level temperature.
An ignored field never causes an error. With Mirai events on,
mirai.session.applied lists each ignored field and the reason. A value that is
wrong, such as a string where a number belongs or a number out of range, rejects
the whole update: you get an error event whose param names the field, and
nothing in that update is applied.
The mirai block
session.mirai carries settings that the OpenAI protocol has no field for. We
validate it strictly: an unknown key returns an error with code
unknown_parameter, so a typo is never silently ignored.
Pipecat 1.8.1’s
SessionProperties drops fields it does not know, so stock
Pipecat cannot send the mirai block. Send your own session.update from a
client that can.
Tools
You can declare function tools insession.update. When the model calls one,
you receive response.output_item.added with a function_call item, then
response.function_call_arguments.delta and .done with the call ID, the tool
name and its arguments. Run the tool, then send:
conversation.item.createwith an item of typefunction_call_output, the samecall_idand your result inoutput.response.create, so the agent continues. The agent also continues on its own once the result arrives, so aresponse.createsent right after the result does not start a second reply.
llm.register_function.
We wait 10 seconds for a result by default (mirai.tool_timeout_ms changes it,
from 1 to 30 seconds). After that, the model gets a tool_timeout result that
asks it to tell the caller the request could not be completed, so the caller is
never left in silence.
With agent_id, the agent’s own tools keep running on our side,
as on any call. A tool you declare with the same name as one of the agent’s
replaces it for the session.
Interruptions
When our turn detection hears the caller start speaking over the agent, you receiveinput_audio_buffer.speech_started. Every reply still in progress ends
with response.done and status cancelled (status_details.reason is
turn_detected). Stop playback, then send conversation.item.truncate with
audio_end_ms, how much of the reply the caller actually heard. We cut the
agent’s line in the conversation to those words, so the model does not believe
it said things the caller never heard. Pipecat does this for you.
response.cancel stops a reply the same way, with reason client_cancelled.
Turn metrics
Everyresponse.done carries the turn’s numbers in response.metadata. The
values are strings, as the OpenAI protocol requires, so any Realtime client
accepts the event:
A time that was not measured for a turn is left out. In the example the model’s
time is missing, and speech recognition and speech synthesis both ran on their
fallbacks.
Stock Pipecat 1.8.1 does not keep
response.metadata when it parses
response.done. To use these numbers there, read the raw event, or turn on
Mirai events in a client that can handle them.
Mirai events
Two events are ours rather than OpenAI’s. We send them only when you ask, by addingmirai_events=1 to the session URL, or by sending a
mirai block in session.update:
mirai.session.applied is sent when the session starts and after every
session.update. It shows what is actually running: the voice, the formats,
the transcription language, turn_detection ("mirai" or "client"),
create_response, interrupt_response, max_output_tokens, the model settings
in llm, tool_timeout_ms, your tools and the agent’s (server_tools), the
providers serving each step, and the length of the prompt in
instructions_chars. Beside it, changed lists the settings that update
changed (["*"] when the session starts), and ignored lists each field we did
not apply, with the reason. Abridged:
mirai.turn.metrics is sent just before each response.done, with the same
numbers as its metadata as JSON numbers, plus the switches between providers
during the turn:
null. Each entry in failovers is
{"kind", "from", "to", "reason"}, for example
{"kind": "tts", "from": "mira", "to": "sarvam", "reason": "no_audio"}.
Fallback
Every session runs Mirai’s own speech recognition, model and speech synthesis. If one of them is unavailable during a session, that step moves to a third-party fallback and the conversation continues:- Speech recognition and speech synthesis fall back to Sarvam.
- The language model falls back to third-party models.
response.done, so you can tell turns that used a
fallback apart and compare their latency.
Billing
A session is billed per minute at your workspace’s tier rate, exactly like a browser call:- Billing starts when the session’s audio goes live and stops when the socket closes.
- Minutes are billed in 10-second blocks, rounded up, with a one-block minimum.
- There is no charge per token, per request or per tool call.
- A session is a call on the
realtimechannel: it appears in your call list and is charged to your wallet like any other call.
Errors
If a session cannot start, the WebSocket upgrade fails with an HTTP status and an OpenAI-style error body:429 Too Many Requests
Clients that authenticate with the subprotocol cannot read an HTTP status, so
they get the upgrade, one
error event with the same code, and a close code of
4000 plus the status (for example 4429).
During a session, errors arrive as error events:
Limits
- Maximum duration. A session without
agent_idruns for up to 300 seconds. Withagent_id, the agent’s own limit applies. Setmax_duration_secsin the URL to change it, from 30 to 1800 seconds. At the limit you get anerrorwith codesession_expired, and the socket closes. - Concurrency. Sessions count towards your workspace’s concurrent calls
(see limits). A session over the limit is refused with
429 concurrency_limitandRetry-After: 5. It does not queue. - Setup time. If setting up a session takes longer than about 8 seconds,
for example because an agent’s pre-call tools are slow, the upgrade fails with
504 start_timeout. - First update. We wait 1.5 seconds after
session.createdfor your firstsession.update. - Message size. A single WebSocket message can be up to 16 MB.
- Idle connections. We ping your client every 30 seconds. A client that stops answering pings is disconnected, and the session ends.
- Servers only. Browser and mobile clients need short-lived credentials, which are not available yet.
Pipecat
Use Mirai voices in a Pipecat pipeline that keeps its own STT and LLM.
Browser calls
Run one of your agents in a user’s browser, without your own server in the audio path.