Skip to main content
In this cookbook you build a voice interview that runs in the candidate’s browser. The same pattern fits any timed, spoken session: screening interviews, mock interviews, oral assessments, language practice or structured feedback calls.

What you’ll build

  • A candidate opens your page, allows the microphone and talks to an AI interviewer that greets them by name.
  • The page shows live captions for both speakers and whether the interviewer is listening, thinking or speaking.
  • The interview keeps time on its own. At 8:00 the interviewer stops starting new topics, at 9:30 it thanks the candidate and says goodbye, and 10:00 is a hard stop. After 60 seconds without an answer it closes politely.
  • A moderator on your side can nudge the interviewer or end the interview gracefully while it runs.
  • When it ends, your backend receives one webhook with the full transcript and a structured score, tagged with your own interview and candidate IDs.
  • No audio is stored. The transcript is the record.
You need: a workspace API key (sk_live_…) and webhook secret (whsec_…), a backend in Python 3.10+ or Node.js 18+, and an HTTPS URL that can receive webhooks (a tunnel to your laptop works while you build).
Keep the API key on your server. The browser only ever receives two per-call links, which work for that one interview and nothing else.

How it works


Step 1: Create the interviewer

An agent holds everything that is the same for every interview of one kind: the interviewer’s instructions, voice and language, how it ends a call, and how the finished conversation is scored. Everything about one candidate goes on the call in step 2.

Write the instructions

The agent’s system_prompt holds up to 8,000 characters. An interview guide fits comfortably if it contains only what the interviewer needs during the conversation. Move everything that is only needed afterwards, such as the scoring rubric, into analysis, which has its own instructions. This layout works well for voice interviews:
system_prompt
The {{placeholders}} are filled per interview from the call’s variables (up to 32 variables, 512 characters each). One agent then serves every candidate for that role type. See the prompting guide for more on writing for voice.
Keep one agent per interview type (for example “Sales screening” and “Support screening”) rather than one agent per candidate. Changing a prompt is then a single edit, and every interview of that type stays comparable.

Score the interview automatically

Turn on analysis and the agent scores each finished interview against your rubric. The result arrives with the transcript in step 5.

Create the agent

interviewer.json
Keep a person in the loop. The score is a summary to help your team review interviews faster, not a hiring decision. The schema above deliberately offers advance or human_review, never reject. Automated decisions about people are regulated in many places, including under the GDPR and India’s DPDP Act.

Step 2: Start an interview from your backend

When a candidate opens the interview page, your backend creates the call with everything specific to this candidate, stores the call ID beside its own interview record, and gives the page two links. Send an Idempotency-Key built from your interview ID. If the request is retried, you get the same call back instead of a second one.
server.py
What to know about the two links:
  • ws_url is the audio socket. It works once and must be opened within 5 minutes. If the candidate refreshes the page, create a new interview.
  • events_url is the live caption stream. It can be reopened for the whole interview and for one hour after it ends.
  • If the page never connects, the call ends as failed with ended_reason: "no-media" and is not billed.

Step 3: Build the interview page

The page asks for the microphone, starts the interview, plays the interviewer, and shows captions, the interviewer’s state and a countdown. The audio format is described in Browser calls → Audio socket.
interview.html
Choices you can change:
  • Barge-in. The candidate can talk over the interviewer, which feels natural. For a stricter format, send silence while the interviewer speaks: track agent.state and fill the microphone frame with zeros while it is speaking.
  • Captions. Each caption appears when its turn is complete. An interrupted interviewer turn shows only the words the candidate actually heard.
  • Leaving. Closing the audio socket ends the interview immediately. To end with a spoken goodbye instead, call your backend’s end route from step 4.

Step 4: Let a moderator steer the interview

A moderator, such as a recruiter watching the captions, can change the interviewer’s direction or end the interview gracefully while it runs. Both go through live control from your backend.
  • An instruction is added to the interviewer’s context from its next reply and stays for the rest of the interview. It is never read aloud and does not appear in the transcript.
  • A close speaks a goodbye line and ends the interview as completed, with ended_reason: "api-ended-call".
Useful instructions for interviews:
server.py (continued)
A moderator view can show the same captions as the candidate. Give your backend the call’s events_url, or read the stream server-side with your API key; see live events.

Step 5: Receive the transcript and the score

Your webhook receives two kinds of events for each interview:
  1. A terminal event as soon as the interview ends: call.completed, call.failed or call.aborted. Mark the interview as over.
  2. call.processed, after the transcript has settled (within about two minutes) and the score is ready. It carries the transcript in data.transcript and the score in data.call.analysis.
Both carry your metadata, so you find the interview without a lookup table. Verify the signature on every delivery and ignore repeats by event id: see Webhooks → Signature verification.
server.py (continued)
What the parts of call.processed mean for an interview: If your webhook was down, nothing is lost: deliveries are retried, and you can always read GET /v2/calls/{id} (it reports transcript_status) and GET /v2/calls/{id}/transcript.

Test it

Short timers show every behaviour in about two minutes. Use them while you build, then switch back to the 10-minute values.
Test values for POST /v2/calls
Run three interviews:
  1. Talk for the full two minutes. At 1:00 the interviewer stops starting new topics. At 1:40 it says the close message and the page shows “Time’s up”. Your webhook receives call.completed, then call.processed with ended_reason: "time-limit-close", the transcript and a score.
  2. Say one sentence, then stay silent. About 30 seconds later the interviewer says the silence message and the interview ends as user-silence-close.
  3. Steer and end from the moderator routes. An instruction returns applied and the next reply follows it. The end route makes the interviewer say goodbye; the interview ends as api-ended-call.

Privacy and compliance

  • Tell candidates up front that they are talking to an AI interviewer, that the conversation is transcribed, and how the transcript and score will be used. Get their consent before the interview starts.
  • No audio is kept with recording_enabled: false: no recording is stored and GET /v2/calls/{id}/recording returns 404. The transcript is kept; see data retention. Ask help@miraiminds.co to turn recording off for your whole workspace.
  • Keep metadata opaque. Put IDs in metadata, never names, contact details or interview answers.
  • Keep a person in the loop for hiring decisions. Use the score to sort and review, not to reject automatically.
  • Protect the two links. ws_url and events_url grant access to one interview. Send them only to that candidate’s page over HTTPS, and keep them out of logs and analytics. Your API key never leaves your server.
  • Treat captions as untrusted text. Render them with textContent, never as HTML.

Production checklist

  • One agent per interview type, with a Timekeeping section and an analysis rubric.
  • metadata carries your interview and candidate IDs; your database stores the returned call id.
  • Every create uses an Idempotency-Key built from your interview ID.
  • max_duration_secs, timing, recording_enabled: false, final_results: true and webhook_url are set on every interview.
  • Moderator routes check that the caller is your staff and use idempotency keys.
  • The webhook verifies signatures, ignores repeated event IDs, answers quickly and processes call.processed in the background.
  • Your interview status is driven by the terminal event, and your transcript processing by call.processed.
  • Candidates see an AI disclosure and give consent before the interview starts.
  • You ran the three test interviews above on the tier you will use in production.

Browser calls

Audio format, live events, live control and timing reference.

Webhooks

Event types, signature verification, retries and final results.

Agents

Every agent field, voices and ending a call.

Limits

Prompt, metadata, timing and data-retention limits.