Skip to main content
A browser call runs an agent with your user’s microphone and speakers instead of a phone line. It is the same call as a phone call — same agent, same tier, same billing from the moment audio goes live, same webhooks — created with channel: "web" and no to. Your backend keeps the API key. It creates the call and hands the page two short-lived links: one for audio, one for live events. The page never sees a secret key.
Building a timed, spoken session such as an interview? The timed voice interviews cookbook puts this page together end to end, with a backend, a page and scoring.
Base URL https://sandbox.voice.miraiminds.co. API requests use Authorization: Bearer sk_live_….

How it fits together

  1. Your page asks your backend to start a session.
  2. Your backend calls POST /v2/calls with channel: "web", stores the returned call id beside its own session, and returns ws_url and events_url to the page.
  3. The page opens the audio socket at ws_url and the event stream at events_url. The call moves from queued to in_progress, and call.started is sent to your webhook.
  4. While the call runs, the page shows each turn and the agent’s state. Your backend can send an instruction or close the call.
  5. The call ends. The page receives call.ended on the event stream. Your webhook receives one terminal event, then — if you opted in — call.processed with the final transcript.
A browser call never rings anything, so it skips dialing: it stays queued until the page connects its audio, then goes straight to in_progress.

Create the call

Create the call from your backend, with the same request as a phone call, plus channel: "web" and without to. Sending to with channel: "web" returns 400 invalid_request: no number is ever dialled. The fields below matter most for browser calls. The full list, with limits, is in Calls. Send an Idempotency-Key derived from your own session, so a retried request never creates a second call.
202 Accepted
Both URLs are absolute and ready to use. Pass them to the page as they are — do not build them yourself.
The two links are credentials for this one call. Anyone holding ws_url can speak to your agent; anyone holding events_url can read the conversation. Send them only to the page running this call, over HTTPS, and keep them out of logs and analytics. Never send your sk_live_ key to a browser.
What happens around the links:
  • The page never connects. If no audio connects before ws_url expires, the call ends as failed with ended_reason: "no-media" and is not billed.
  • You retry the create request. With the same Idempotency-Key, you get the same call. The replay includes ws_url only while that link is still unused.
  • The audio drops mid-call. ws_url cannot be used twice, so a dropped audio socket cannot be reconnected. Create a new call to continue.

Audio socket

Open ws_url as a WebSocket from the page. Audio travels as binary messages in both directions; the server also sends one kind of text message.
  • Send small frames, continuously. 20 ms per message (960 samples, 1,920 bytes) works well.
  • To mute, send silence. Keep sending frames filled with zeros instead of stopping. The agent then hears a quiet line, not a dropped one.
  • Keep echo cancellation on (echoCancellation: true), so a user on speakers does not feed the agent its own voice.
  • Ignore anything else. Text messages you send are ignored. Skip any text message from the server that you do not recognise.
  • Closing the socket ends the conversation. Do it when the user leaves the page.
Browsers allow microphone access only on secure pages: serve the page over HTTPS (localhost is fine while you build). The socket accepts any origin that presents a valid token, so the page can live on your own domain.
The browser does not tell your code why a WebSocket or EventSource was refused. If the audio socket closes before it opens, the link has expired or was already used. Create a new call.

Live events

Open events_url with EventSource. It is a Server-Sent Events stream: the page receives the agent’s state and each completed turn as the conversation happens.

Framing

Each event is one SSE message with three lines:
  • id is the position in the stream. Treat it as opaque: it is what a reconnect resumes from.
  • event is the event type. Because every event is named, listen with addEventListener(type, …). onmessage receives nothing.
  • data is a JSON envelope:
Treat the envelope as additive: new fields and new event types can appear. Ignore what you do not recognise.

agent.state

The page’s own playback buffer can add a short delay between speaking and the moment sound comes out of the speaker.

transcript.turn

One event per completed turn, for both speakers.
Text arrives once per completed turn: the user’s after they finish speaking, the agent’s after its turn ends. There are no word-by-word captions yet.

call.ended

Always the last event. The stream closes after it.
call.ended is sent even when the conversation ends abnormally. Use it to end the session in the page. The webhook is the record your backend acts on.

Reconnecting

  • Automatic resume. EventSource reconnects on its own and sends the Last-Event-ID header. The stream continues after that event, so no turn is missed or repeated.
  • Manual resume. Add after=<id> to events_url to start after a given event, for example after a page reload.
  • One-hour buffer. Events are kept for one hour after the call ends. Opening the stream for a finished call replays what happened, then closes.
  • Close it yourself. Call events.close() when you receive call.ended. Otherwise EventSource keeps reconnecting to a stream that has nothing left to send.
  • Keep-alives. The server sends an SSE comment every 15 seconds. EventSource ignores these; they keep proxies from closing an idle stream.
The stream allows every origin (Access-Control-Allow-Origin: *) and uses no cookies: the token in events_url is the credential. Your backend can read the same stream with Authorization: Bearer sk_live_… instead of the token:

Live control

Steer or close a browser call while it is in progress. Send it from your backend with your API key. If the user starts the action in the page, the page asks your backend.
  • instruction adds your text to the agent’s context. The agent follows it from its next reply for the rest of the call. The text is not read aloud.
  • close speaks message, then ends the call. Without message, the agent speaks its end_call.message, or else a short goodbye of its own. The call ends even if the agent would not have chosen to end it. It ends completed with ended_reason: "api-ended-call".
200 OK
The response is 200 when the live call answered within 3 seconds, with applied or rejected. Otherwise it is 202 with status: "pending". Check a pending control with GET:
Idempotency. Send an Idempotency-Key. A repeat with the same key returns the same control record and is never delivered twice, so a retry cannot make the agent say goodbye twice.

Close or abort

Both end a live browser call. They are different tools. Use close for a graceful ending the user hears. Use abort when the conversation must stop now. If a close, an abort, the user leaving, and a timer all happen at once, the call still ends exactly once: one terminal status, one terminal webhook, and at most one goodbye.

Timing policy

timing on POST /v2/calls lets the call end gracefully instead of being cut off at its hard cap. It works on browser and phone calls. Every field is optional. max_duration_secs here is the call’s effective cap: the value on the request, or else the agent’s. Set each timer together with its message. How the clock works:
  • It starts when audio goes live: when the page connects, or when the phone is answered. Time spent queued does not count.
  • The hard cap still applies. If a timed close cannot finish, the call still ends at max_duration_secs (status: timeout, ended_reason: "exceeded-max-duration").
  • Only the user resets the silence timer. The agent’s own speech does not count as a user response.
  • Placeholders work. {{variable}} placeholders in the messages and the instruction are filled from the call’s variables.
For max_duration_secs: 600 with the example above: Invalid timing returns 400 invalid_request.

Browser example

A complete page: microphone capture with an AudioWorklet, PCM16 over the audio socket, playback with interruption, and a live transcript for both speakers. It assumes the backend routes in the next section.
index.html
Notes on the sample:
  • Barge-in is on by default. The user can talk over the agent; the server sends clear and the speaker worklet drops what it had queued. Uncomment the muted line to turn barge-in off. The microphone then sends silence while the agent speaks, so the user cannot interrupt.
  • Sample rate. Most browsers convert the microphone into a 48 kHz AudioContext for you. If one does not, capture at the device’s rate and resample to 48,000 Hz before sending.
  • Playback smoothing. The speaker worklet plays audio as soon as it arrives. On unreliable networks, buffer 100–200 ms before starting each reply.

Backend example

Three routes in Node.js with Express: start a session, end it gracefully, and receive webhooks. db, queue, requireUser, verifySignature and alreadyProcessed stand in for your own code. The last two are on the webhooks page.
server.js

Ending and results

A browser call sends the same webhooks as a phone call, in this order:
  1. call.queued — only if enabled for your workspace.
  2. call.started — the page connected its audio.
  3. Exactly one terminal event:
    • call.completed — the conversation finished: the user left, the agent ended it, or a close, timed close or silence close ended it.
    • call.failed — it never connected (status: failed, ended_reason: "no-media", not billed), it hit the hard cap (status: timeout, billed), or it failed on our side (not billed).
    • call.aborted — you called POST /abort.
  4. call.processed — when you set final_results: true (or the agent has analysis or post-call tools). With final_results, it carries the finished transcript.
Which event to act on: call.processed is not the “session over” signal. It arrives after the terminal event: the transcript settles within about 2 minutes of the call ending, plus the time analysis and post-call tools take when the agent has them. Mark the session over on the terminal event, and start your transcript processing on call.processed. data.transcript.status is one of: Deliveries are at least once, and order is not guaranteed. Deduplicate on the event id — retries keep it — and order your own state by data.call.status, not by arrival. data.call.metadata carries your session_id on every event, so you can match an event even before your own create request has finished writing the call ID.

Privacy and security

  • Keys stay on your server. The page gets two links scoped to one call.
  • Keep metadata opaque. metadata is stored with the call and sent on every webhook. Put IDs in it, not secrets or personal answers.
  • Turn off stored audio when you do not need it. recording_enabled: false keeps no recording. Live audio is still processed during the call, and the transcript is still kept.
  • Tell users. Say in the opening line that they are speaking with an AI and whether the conversation is recorded or transcribed. Collect the consent your jurisdiction requires (for example under GDPR or India’s DPDP Act) before you start the call.

FAQ

POST /v2/call/web is the v1 endpoint. Use POST /v2/calls with channel: "web". You get a call ID, an audio link, a live event link, live control, timing, metadata on every webhook and the final transcript.
An agent’s system_prompt holds up to 8,000 characters. There is no per-call system prompt. Per call, you can set the opening line with first_message (up to 500 characters), pass data into the prompt with variables, and direct the agent mid-call with an instruction.
Yes. Set max_duration_secs: 600 on the call. The accepted range is on Limits. Add a timing policy so the agent wraps up and says goodbye before the cap.
No. Each call’s events go only to the webhook_url it was created with. A call without one sends no webhooks. Every delivery is signed with your workspace’s webhook secret.
Yes. Create the call with recording_enabled: false. No audio is stored, recording_available is false, and GET /v2/calls/{id}/recording returns 404. The transcript, live events and billing are unchanged. To turn recording off for your whole workspace, ask support. After that, a request with recording_enabled: true returns 403 recording_disabled.
Transcripts follow the data retention policy. If you need a different retention period, ask support.
Not yet. The event stream sends each turn when it is complete. For an interrupted agent turn, the text is what the user actually heard.
No. metadata is for your systems only. Anything the agent should know goes in variables.