Skip to main content
The Realtime API speaks the OpenAI Realtime protocol. If your bot already uses Pipecat’s OpenAIRealtimeLLMService, or any other client for that protocol, you change the API key and the URL and keep everything else:
Each session runs Mirai’s speech recognition, turn detection, language model and speech synthesis on one voice worker, next to each other. Your process sends the caller’s audio up and plays the agent’s audio back. Compared with calling our speech-to-text, LLM and text-to-speech APIs one by one, a turn makes one round trip to your server instead of three. A session is a call. It has a call ID, it shows up in your call list, it sends webhooks and it is billed per minute, like a browser call.

Connect

Use the API address shown under Developers in the console: prod.voice once your company has gone live, sandbox.voice.miraiminds.co with a sandbox key. A malformed parameter fails the upgrade with 400 invalid_request, and error.param names the parameter. Pipecat builds its URL as base_url + "?model=…". A base_url that already carries a query, such as …/v2/realtime?agent_id=agt_…, therefore reaches us as ?agent_id=agt_…?model=…. That is fine: we read your parameters correctly.

Authentication

Send your workspace API key in the Authorization header, as every OpenAI Realtime client does:
Clients that cannot set headers can send the key as a WebSocket subprotocol instead, openai-insecure-api-key.sk_live_…, alongside realtime. We answer with the realtime subprotocol.
The Realtime API is for servers. Never put your API key in a web page or a mobile app: anyone who opens the page can read it. Short-lived browser credentials are not available yet.

With Pipecat

Stock Pipecat 1.8.1 works unchanged. Here is a complete microphone-and-speaker agent. Save it as pipecat_realtime.py, put on headphones and run it with uv, which installs Pipecat for you. The microphone needs PortAudio (brew install portaudio on macOS).
Two things to set in your own pipeline:
  • Send and play 16-bit PCM at 24 kHz. Set audio_in_sample_rate=24000 and audio_out_sample_rate=24000 in PipelineParams.
  • Remove your own STT, LLM and TTS services. The Realtime service replaces all three, and your transport and context aggregators stay as they are.
The agent never greets on its own in a Realtime session. Queue an LLMRunFrame (or send response.create) when you want it to speak first.

Sessions

  1. You open the socket. We check your key and balance and start a call on the realtime channel.
  2. We send session.created.
  3. You send session.update with your settings. We apply them and answer session.updated. We wait up to 1.5 seconds after session.created for this first update. If it does not arrive in time, the session starts with the agent’s settings (or the defaults), and your update is applied when it comes, apart from the language and audio formats.
  4. You stream audio with input_audio_buffer.append. We detect when the caller has finished, transcribe, run the model and stream the reply as audio. Audio you send before the session is ready is kept, up to 2 seconds of it.
  5. You close the socket, or the session reaches its maximum duration. The call ends and is billed.
You can send session.update again at any point. Most changes apply straight away; session settings lists the exceptions. Every event travels as a JSON text frame, audio included (base64). Binary frames are ignored.

Events

The events are the OpenAI Realtime GA events. These are the ones we handle. You send output_audio_buffer.clear and transcription_session.update are accepted and ignored. Any other event type returns an error with code invalid_event. We send Not supported yet
  • rate_limits.updated.
  • Interim transcripts. Transcription deltas carry final segments only, so a client that joins the deltas never prints a word twice.
  • Text-only replies. Every reply is spoken, so there is no response.output_text.delta.
  • Per-response settings in response.create.
  • Image and audio content in conversation.item.create. Only text parts are read.

Session settings

“First update only” means the session.update that arrives within 1.5 seconds of session.created. A later change to the language or an audio format is ignored, because both are fixed once the session has started. Fields we do not know are ignored. Beta-shaped fields are accepted too: input_audio_format and output_audio_format (pcm16, g711_ulaw, g711_alaw), input_audio_transcription, turn_detection, voice, speed, modalities, max_response_output_tokens and a top-level temperature. An ignored field never causes an error. With Mirai events on, mirai.session.applied lists each ignored field and the reason. A value that is wrong, such as a string where a number belongs or a number out of range, rejects the whole update: you get an error event whose param names the field, and nothing in that update is applied.

The mirai block

session.mirai carries settings that the OpenAI protocol has no field for. We validate it strictly: an unknown key returns an error with code unknown_parameter, so a typo is never silently ignored.
Pipecat 1.8.1’s SessionProperties drops fields it does not know, so stock Pipecat cannot send the mirai block. Send your own session.update from a client that can.

Tools

You can declare function tools in session.update. When the model calls one, you receive response.output_item.added with a function_call item, then response.function_call_arguments.delta and .done with the call ID, the tool name and its arguments. Run the tool, then send:
  1. conversation.item.create with an item of type function_call_output, the same call_id and your result in output.
  2. response.create, so the agent continues. The agent also continues on its own once the result arrives, so a response.create sent right after the result does not start a second reply.
Pipecat does both for functions registered with llm.register_function. We wait 10 seconds for a result by default (mirai.tool_timeout_ms changes it, from 1 to 30 seconds). After that, the model gets a tool_timeout result that asks it to tell the caller the request could not be completed, so the caller is never left in silence. With agent_id, the agent’s own tools keep running on our side, as on any call. A tool you declare with the same name as one of the agent’s replaces it for the session.

Interruptions

When our turn detection hears the caller start speaking over the agent, you receive input_audio_buffer.speech_started. Every reply still in progress ends with response.done and status cancelled (status_details.reason is turn_detected). Stop playback, then send conversation.item.truncate with audio_end_ms, how much of the reply the caller actually heard. We cut the agent’s line in the conversation to those words, so the model does not believe it said things the caller never heard. Pipecat does this for you. response.cancel stops a reply the same way, with reason client_cancelled.

Turn metrics

Every response.done carries the turn’s numbers in response.metadata. The values are strings, as the OpenAI protocol requires, so any Realtime client accepts the event:
A time that was not measured for a turn is left out. In the example the model’s time is missing, and speech recognition and speech synthesis both ran on their fallbacks. Stock Pipecat 1.8.1 does not keep response.metadata when it parses response.done. To use these numbers there, read the raw event, or turn on Mirai events in a client that can handle them.

Mirai events

Two events are ours rather than OpenAI’s. We send them only when you ask, by adding mirai_events=1 to the session URL, or by sending a mirai block in session.update:
Do not set mirai_events with stock Pipecat. Pipecat 1.8.1’s Realtime service raises an error on any event type it does not know, and that stops it reading the socket, so the session goes silent. Turn it on only in a client that ignores or handles event types it does not recognise.
mirai.session.applied is sent when the session starts and after every session.update. It shows what is actually running: the voice, the formats, the transcription language, turn_detection ("mirai" or "client"), create_response, interrupt_response, max_output_tokens, the model settings in llm, tool_timeout_ms, your tools and the agent’s (server_tools), the providers serving each step, and the length of the prompt in instructions_chars. Beside it, changed lists the settings that update changed (["*"] when the session starts), and ignored lists each field we did not apply, with the reason. Abridged:
mirai.turn.metrics is sent just before each response.done, with the same numbers as its metadata as JSON numbers, plus the switches between providers during the turn:
A time that was not measured is null. Each entry in failovers is {"kind", "from", "to", "reason"}, for example {"kind": "tts", "from": "mira", "to": "sarvam", "reason": "no_audio"}.

Fallback

Every session runs Mirai’s own speech recognition, model and speech synthesis. If one of them is unavailable during a session, that step moves to a third-party fallback and the conversation continues:
  • Speech recognition and speech synthesis fall back to Sarvam.
  • The language model falls back to third-party models.
You pay the same per-minute rate either way. The providers that served each turn are in response.done, so you can tell turns that used a fallback apart and compare their latency.

Billing

A session is billed per minute at your workspace’s tier rate, exactly like a browser call:
  • Billing starts when the session’s audio goes live and stops when the socket closes.
  • Minutes are billed in 10-second blocks, rounded up, with a one-block minimum.
  • There is no charge per token, per request or per tool call.
  • A session is a call on the realtime channel: it appears in your call list and is charged to your wallet like any other call.
A session needs enough balance to start, and ends if your balance runs out while it is open.

Errors

If a session cannot start, the WebSocket upgrade fails with an HTTP status and an OpenAI-style error body:
429 Too Many Requests
Clients that authenticate with the subprotocol cannot read an HTTP status, so they get the upgrade, one error event with the same code, and a close code of 4000 plus the status (for example 4429). During a session, errors arrive as error events:
Stock Pipecat 1.8.1 treats every error event as fatal and stops reading the socket, except conversation_already_has_active_response, response_cancel_not_active and item_retrieve_invalid_item_id. Check the values you send in session.update before you go live: a rejected update ends a Pipecat session.

Limits

  • Maximum duration. A session without agent_id runs for up to 300 seconds. With agent_id, the agent’s own limit applies. Set max_duration_secs in the URL to change it, from 30 to 1800 seconds. At the limit you get an error with code session_expired, and the socket closes.
  • Concurrency. Sessions count towards your workspace’s concurrent calls (see limits). A session over the limit is refused with 429 concurrency_limit and Retry-After: 5. It does not queue.
  • Setup time. If setting up a session takes longer than about 8 seconds, for example because an agent’s pre-call tools are slow, the upgrade fails with 504 start_timeout.
  • First update. We wait 1.5 seconds after session.created for your first session.update.
  • Message size. A single WebSocket message can be up to 16 MB.
  • Idle connections. We ping your client every 30 seconds. A client that stops answering pings is disconnected, and the session ends.
  • Servers only. Browser and mobile clients need short-lived credentials, which are not available yet.

Pipecat

Use Mirai voices in a Pipecat pipeline that keeps its own STT and LLM.

Browser calls

Run one of your agents in a user’s browser, without your own server in the audio path.