Skip to main content
Mira Transcribe turns a recording into text: what was said, who said it and when. You send a file or a link and get a job ID straight away. A draft is readable within seconds, and the finished transcript replaces it when it is ready. A two-minute call usually takes about 30 seconds. Base URL https://sandbox.voice.miraiminds.co. API requests use Authorization: Bearer sk_live_…, the same workspace key as the rest of the v2 API.

When to use it

Turning example calls into a node-based agent has its own transcription step, priced separately. See Draft a flow from call recordings.

How it works

Two speech recognition passes listen to the whole recording. A language model then compares them and writes one transcript, using any context and vocabulary you sent to settle names and terms. Speaker separation listens to the voices themselves to work out who spoke when, and gives each voice its own ID. While the language model works, the job carries a draft: the stronger of the two passes, with timestamps and no speakers. In the draft, English is often written out in Devanagari. The finished transcript writes Hindi in Devanagari and English in Latin script, the way the speakers switched between them. Common English loanwords can stay in Devanagari. Nothing is translated.
A job moves through queued, decoding and merging to completed or failed. The reference describes each status.

Try it in the console

1

Open the Playground

Sign in to the console, open Speech to text and choose Offline.
2

Add recordings

Choose one recording or several: WAV, MP3 or M4A, up to 45 MB each. Add a line of context and any names or terms to spell right, then click Transcribe.
3

Open the transcript

Each recording gets a row with its length, a clock and its stage. Open it as soon as the draft appears. The finished transcript takes its place in the player on its own.
The player shows the recording as a timeline with a lane for each speaker. The transcript scrolls with the audio, and clicking a line plays from there. You can rename, recolour and merge speakers, move a line to another speaker, fix words, and split or join lines, with undo. Exports are plain text, SRT and WebVTT subtitles, and JSON. Edits are kept in your browser, which also keeps a copy of your last few uploads for up to a week so you can play them back. Logs lists every recording your workspace sent in the last 30 days, from the console or the API, with what each one cost. API has the requests on this page ready to copy, and can create a key for you. Jobs from the console are charged to the workspace wallet in the same way as API jobs.

Use the API

Send the recording as a multipart upload:
The answer is 202 Accepted with the job, whose status is queued. Read it every few seconds until status is completed or failed:
A recording over 45 MB goes by link instead: send JSON with an https url (up to 200 MB), such as a pre-signed link to your storage. The Transcriptions reference has every field, response and error, with Python and Node.js. To build a review screen on top of this, follow the transcript player cookbook.

Speakers and timestamps

A finished job has one segment for each turn:
start and end are seconds from the start of the recording. speaker_id is the voice: spk0, spk1 and so on, numbered in the order the voices first speak. Group lines by it. speaker is a role, Agent, Customer or Other, worked out from what each voice says. It is a best guess and can be wrong, so show it as a first label that a person can change. It is empty when the role could not be told. diarization.speakers lists each voice once, with its role and how many seconds it talked. Send words: true for the time of every word. Word timings are added when they line up with the audio, which today means Hindi recordings; Gujarati and English recordings come back without them. diarization.words.applied tells you whether they were added. They need speaker separation, which is on unless you send diarize: false.

Languages

Mira Transcribe is built for phone calls in Hindi and Hinglish. language defaults to auto, which detects the main language of the call. You can set it to hi, gu, en, mr, pa, bn or ur when you know the main language and detection gets it wrong. Speech in another language inside the call is kept as it was spoken.

Getting names and terms right

vocabulary takes names, brands, products, places and codes, spelled the way you want them back. Send up to 200 terms of up to 100 characters each, as one string separated by commas or new lines, or as a JSON array. A term can’t contain a comma, and repeats are dropped. context is a sentence or two about the call: who is calling whom, and why. It can be up to 2,000 characters. Both affect spelling and script. We sent the same 23-second test call twice. With an unrelated context, “Mira University” and “Aadhaar” came back as “मीरा यूनिवर्सिटी” and “आधार”. With the context and vocabulary in the example above, both came back as written. Use the same context and vocabulary for every call of one kind.

Pricing

Transcription costs ₹18 per hour of audio on every plan. Prices exclude GST.
  • A job is charged by the length of the recording, to the millisecond, rounded up to the paisa once. A 23-second call costs ₹0.12 and a five-minute call ₹1.50.
  • The charge is taken from your wallet once, when the job completes. It is taken whether or not you read the result.
  • A degraded job is free. So is a job that fails for any reason.
  • The rate is fixed on each job when it is accepted, and cost on the job says what it was charged.
  • Starting a job needs a wallet balance above zero, and nothing is held while it runs. With an empty wallet the request is refused with 402 insufficient_balance.
Each charge is one wallet transaction with kind: "transcription", the job ID as reference and the billed seconds as quantity. The billing statement in the console shows transcription under Charges, with the number of recordings and minutes of audio.

Degraded jobs

When the language model step is unavailable, the job still completes with the stronger recognition pass and degraded: true. The words are there with their timestamps, but there are no roles and the text has had less clean-up. Speaker IDs can still be present. A degraded job is not charged. If you need the roles, send the recording again; the new job is charged when it completes in full.

Limits

Over a limit, the request is refused with 413 or 429 and nothing is charged. A recording longer than two hours is accepted, then fails with audio_too_long, free. See limits and errors for every code.

Privacy and your data

The API does not keep your audio. An upload passes through a temporary file that is deleted when the request ends, and the transcription service deletes the recording once it has decoded it. A recording link is fetched once and is not stored. The transcript stays with the transcription service, so that GET /v2/transcriptions/{id} can return it. There is no published retention period for transcripts yet, so copy each finished transcript into your own storage. If a transcript is no longer available, its job still reads completed with its cost, but without text and segments. The job record itself (status, file name, size, length and cost) stays with your workspace and appears in Logs. There is no endpoint to delete a transcript; if you need one removed, write to help@miraiminds.co. Recordings and transcripts of calls are personal data under India’s Digital Personal Data Protection Act, and under laws such as the GDPR when they cover people elsewhere. Tell the people on your calls that they are recorded and transcribed, and get consent where the law asks for it. Keep your API key on your server: send recordings from your backend, never from a browser or a mobile app.

Troubleshooting

Transcriptions API

Every field, status, response and error code.

Build a transcript player

Upload, draft, finished transcript, playback, edits and export in your own web app.

Wallet

Balance, transactions and top-ups.

Call transcripts

Transcripts of calls your voice agents made.