What you’ll build
- A reviewer chooses a recording (WAV, MP3 or M4A, up to 45 MB), adds a line of context and a few names to spell right, and clicks Transcribe.
- A few seconds later the page shows a draft. When the job finishes, the transcript with speakers replaces it without a reload.
- Each speaker has a colour, keyed to the voice (
speaker_id). The roles that come back are a first guess, so the reviewer can rename any speaker. - Clicking a line plays the call from there. While the audio plays, the line being spoken is marked and kept in view.
- Double-clicking a line lets the reviewer fix a word. Edits survive a reload.
- The transcript, edits included, exports as plain text, SRT subtitles or JSON.
- Your server keeps the API key, and only returns a transcript to the person who uploaded it.
sk_live_…) and some credit in its wallet,
and a backend in Python 3.10+ or Node.js 18+. This recipe uses no webhooks.
Keep the API key on your server. The page only ever talks to your own routes,
and job IDs are not secrets, so your server checks who owns a job before it
returns one.
How it works
The API keeps no audio, so the page plays the file the reviewer chose. For the 23-second test call, the page goes through these stages:
A longer recording usually takes about 10 seconds plus 15% of its length, so
about 30 seconds for a two-minute call. The draft arrives before that.
Step 1: Send recordings from your server
Your server holds the key and has two routes for the page:
Refusals from the API, such as
402 insufficient_balance, 413 and 429, go
back to the page with their Retry-After header, so the page can say what
happened. The server also serves the page itself from public/.
myapp stands for your own code. current_user (Python) or requireUser
(Node.js) identifies the signed-in user and refuses everyone else with 401.
db.transcripts stores one row per job: create adds the owner and file name
and ignores an ID it already has, get returns the row or nothing, and
save_job stores the finished job.
- cURL
- Python
- Node.js
The two API calls the server makes:
Step 2: Poll from the page
The page asks your server for the job until it iscompleted or failed. It
asks quickly at first, since a short recording is often done in seconds, then
every 4 seconds. On a 429 or 5xx answer it waits for Retry-After and
asks again, since the job itself is unaffected. It gives up only after about a
minute without a usable answer.
public/poll.js
Step 3: Turn a job into a transcript
The page keeps the transcript as a small document: speakers, and lines that point to a speaker by ID. Renaming a speaker is then one change that every line and every export picks up. This module has no page code in it, so you can test it on its own.public/transcript.js
speaker_id, which identifies a voice. The role in
speaker (Agent, Customer or Other) only supplies the first label,
numbered when two voices share a role. A draft has no speakers yet, so all of
it goes in one lane called Speech. Colours follow the order in which voices
first speak, so the first voice is always green.
Step 4: Build the player page
public/index.html
Draft, then the finished transcript
While the job ismerging, show draws the draft greyed out and read-only,
with export turned off. When the job completes, the finished transcript
replaces it in place. The status line shows what the job cost, or that it was
free. A degraded job is marked
as a single-pass transcript without roles.
Play from a line and follow the audio
Clicking a line setsaudio.currentTime to the line’s start and plays. On
every timeupdate, lineAt finds the line being spoken; the page marks it and
scrolls it into view. If the reviewer scrolls, following pauses for six
seconds, so the list doesn’t pull them back while they read ahead.
Rename speakers and fix words
Each speaker’s name is an input. Changing it renames every line of that speaker. Double-clicking a line’s text makes it editable: Enter keeps the change and Escape drops it. Edits are saved in the browser under the job ID, and the transcript text is always set withtextContent, never as HTML.
Export
The three buttons turn the current document, edits included, into a file:Speaker: text lines, SRT cues with the speaker’s name in front of each line,
or JSON with speaker_id, speaker, start, end and text per line.
Open a transcript later
The page’s address ends with the job ID (#trn_…). Opening that address loads
the transcript from your server and asks for the recording, because the API
does not keep it. If you keep recordings in your own storage, set audio.src
to a link your server signs instead.
Run it
Put the files side by side:- Python
- Node.js
http://localhost:3000.
Test it
Download the test call: 23 seconds, two synthetic voices, Hindi and English. Use this context and vocabulary:- Upload it. A draft appears after about 3 seconds. After about 6 seconds it is replaced by four lines that alternate between Agent and Customer, with “Mira University”, “BBA” and “Aadhaar” in Latin script. The status reads “Finished (₹0.12)”.
- Click the third line. Playback jumps to 0:11 and the line is marked. Let it play: the mark moves to the fourth line at 0:20.
- Rename Agent to a person’s name; both of that speaker’s lines change. Double-click the second line, add punctuation and press Enter. Reload the page: both edits are still there.
- Export SRT. The first cue runs
00:00:00,000 --> 00:00:06,570. - Click Transcribe again without choosing the file again. You get the same job ID back, and the wallet is charged once.
- Open the page’s address in a new tab. The transcript loads from your server; choose the file to listen along.
Privacy and compliance
Tell the people on a call that it is recorded, and that a transcript will be made and reviewed, before you record. Get consent where the law asks for it. Call recordings and transcripts are personal data under India’s Digital Personal Data Protection Act. The API does not keep the audio: it deletes the upload once it has been decoded. Your server passes the file through without keeping it, and the player plays the reviewer’s own copy. Your server does keep every finished job, with everything that was said in the call. Treat it like the recording: limit who can open it, decide how long to keep it, and delete it on schedule. The API does not publish a retention period for transcripts; see Privacy and your data. Edits are saved in the reviewer’s browser. On shared computers, save them on your server instead and clear local storage when the reviewer signs out. Words, speakers and roles can be wrong. Don’t base a decision about a person on a transcript that nobody has checked against the audio. The API key stays in your server’s environment, and both routes check who owns a job before they touch it.Production checklist
- The API key is read from the server’s environment, and no route returns it.
- Both routes check sign-in and job ownership.
- Every upload sends an
Idempotency-Keythat includes the user’s ID. - Your server and any proxy in front of it accept uploads up to 45 MB, and larger files go by link.
-
402,413and429are shown to the reviewer in plain words, andRetry-Afteris honoured. - Finished jobs are saved in your storage and served from there. A background task reads jobs that nobody opened, so every transcript reaches your storage.
- Edits are saved where your reviewers need them: on your server if more than one person reviews a call.
- You have a retention period for transcripts and enforce it.
- Callers are told about recording and transcription.
- You ran the test call above and got four lines for ₹0.12.
Related
Transcribe recorded calls
How transcription works, languages, vocabulary, pricing and privacy.
Transcriptions API
Every field, status, response and error code.
Wallet
Balance, transactions and top-ups.
Limits
Request rate and other limits shared by every
/v2 call.