What you’ll build
- A candidate opens your page, allows the microphone and talks to an AI interviewer that greets them by name.
- The page shows live captions for both speakers and whether the interviewer is listening, thinking or speaking.
- The interview keeps time on its own. At 8:00 the interviewer stops starting new topics, at 9:30 it thanks the candidate and says goodbye, and 10:00 is a hard stop. After 60 seconds without an answer it closes politely.
- A moderator on your side can nudge the interviewer or end the interview gracefully while it runs.
- When it ends, your backend receives one webhook with the full transcript and a structured score, tagged with your own interview and candidate IDs.
- No audio is stored. The transcript is the record.
sk_live_…) and webhook secret
(whsec_…), a backend in Python 3.10+ or Node.js 18+, and an HTTPS URL that
can receive webhooks (a tunnel to your laptop works while you build).
Keep the API key on your server. The browser only ever receives two per-call
links, which work for that one interview and nothing else.
How it works
Step 1: Create the interviewer
An agent holds everything that is the same for every interview of one kind: the interviewer’s instructions, voice and language, how it ends a call, and how the finished conversation is scored. Everything about one candidate goes on the call in step 2.Write the instructions
The agent’ssystem_prompt holds up to 8,000 characters. An interview guide
fits comfortably if it contains only what the interviewer needs during the
conversation. Move everything that is only needed afterwards, such as the
scoring rubric, into analysis, which has
its own instructions.
This layout works well for voice interviews:
system_prompt
{{placeholders}} are filled per interview from the call’s variables
(up to 32 variables, 512 characters each). One agent then serves every
candidate for that role type. See the prompting guide for
more on writing for voice.
Score the interview automatically
Turn onanalysis and the agent scores each finished interview against your
rubric. The result arrives with the transcript in step 5.
Create the agent
- cURL
- Python
- Node.js
interviewer.json
Step 2: Start an interview from your backend
When a candidate opens the interview page, your backend creates the call with everything specific to this candidate, stores the call ID beside its own interview record, and gives the page two links.
Send an
Idempotency-Key built from your interview ID. If the request is
retried, you get the same call back instead of a second one.
- Python
- Node.js
server.py
ws_urlis the audio socket. It works once and must be opened within 5 minutes. If the candidate refreshes the page, create a new interview.events_urlis the live caption stream. It can be reopened for the whole interview and for one hour after it ends.- If the page never connects, the call ends as
failedwithended_reason: "no-media"and is not billed.
Step 3: Build the interview page
The page asks for the microphone, starts the interview, plays the interviewer, and shows captions, the interviewer’s state and a countdown. The audio format is described in Browser calls → Audio socket.interview.html
- Barge-in. The candidate can talk over the interviewer, which feels
natural. For a stricter format, send silence while the interviewer speaks:
track
agent.stateand fill the microphone frame with zeros while it isspeaking. - Captions. Each caption appears when its turn is complete. An interrupted interviewer turn shows only the words the candidate actually heard.
- Leaving. Closing the audio socket ends the interview immediately. To end with a spoken goodbye instead, call your backend’s end route from step 4.
Step 4: Let a moderator steer the interview
A moderator, such as a recruiter watching the captions, can change the interviewer’s direction or end the interview gracefully while it runs. Both go through live control from your backend.- An instruction is added to the interviewer’s context from its next reply and stays for the rest of the interview. It is never read aloud and does not appear in the transcript.
- A close speaks a goodbye line and ends the interview as
completed, withended_reason: "api-ended-call".
- Python
- Node.js
server.py (continued)
events_url, or read the stream server-side with your API key; see
live events.
Step 5: Receive the transcript and the score
Your webhook receives two kinds of events for each interview:- A terminal event as soon as the interview ends:
call.completed,call.failedorcall.aborted. Mark the interview as over. call.processed, after the transcript has settled (within about two minutes) and the score is ready. It carries the transcript indata.transcriptand the score indata.call.analysis.
metadata, so you find the interview without a lookup table.
Verify the signature on every delivery and ignore repeats by event id: see
Webhooks → Signature verification.
- Python
- Node.js
server.py (continued)
call.processed mean for an interview:
If your webhook was down, nothing is lost: deliveries are
retried, and you can always read
GET /v2/calls/{id} (it reports transcript_status) and
GET /v2/calls/{id}/transcript.
Test it
Short timers show every behaviour in about two minutes. Use them while you build, then switch back to the 10-minute values.Test values for POST /v2/calls
- Talk for the full two minutes. At 1:00 the interviewer stops starting new
topics. At 1:40 it says the close message and the page shows “Time’s up”.
Your webhook receives
call.completed, thencall.processedwithended_reason: "time-limit-close", the transcript and a score. - Say one sentence, then stay silent. About 30 seconds later the
interviewer says the silence message and the interview ends as
user-silence-close. - Steer and end from the moderator routes. An instruction returns
appliedand the next reply follows it. The end route makes the interviewer say goodbye; the interview ends asapi-ended-call.
Privacy and compliance
- Tell candidates up front that they are talking to an AI interviewer, that the conversation is transcribed, and how the transcript and score will be used. Get their consent before the interview starts.
- No audio is kept with
recording_enabled: false: no recording is stored andGET /v2/calls/{id}/recordingreturns404. The transcript is kept; see data retention. Ask help@miraiminds.co to turn recording off for your whole workspace. - Keep metadata opaque. Put IDs in
metadata, never names, contact details or interview answers. - Keep a person in the loop for hiring decisions. Use the score to sort and review, not to reject automatically.
- Protect the two links.
ws_urlandevents_urlgrant access to one interview. Send them only to that candidate’s page over HTTPS, and keep them out of logs and analytics. Your API key never leaves your server. - Treat captions as untrusted text. Render them with
textContent, never as HTML.
Production checklist
- One agent per interview type, with a Timekeeping section and an
analysisrubric. -
metadatacarries your interview and candidate IDs; your database stores the returned callid. - Every create uses an
Idempotency-Keybuilt from your interview ID. -
max_duration_secs,timing,recording_enabled: false,final_results: trueandwebhook_urlare set on every interview. - Moderator routes check that the caller is your staff and use idempotency keys.
- The webhook verifies signatures, ignores repeated event IDs, answers
quickly and processes
call.processedin the background. - Your interview status is driven by the terminal event, and your
transcript processing by
call.processed. - Candidates see an AI disclosure and give consent before the interview starts.
- You ran the three test interviews above on the tier you will use in production.
Related
Browser calls
Audio format, live events, live control and timing reference.
Webhooks
Event types, signature verification, retries and final results.
Agents
Every agent field, voices and ending a call.
Limits
Prompt, metadata, timing and data-retention limits.