Audio to text
Transcribe audio to text: upload a recording, get timestamped text
Transcribe audio to text by uploading the recording: drop an MP3, M4A, WAV, FLAC or OGG file below, and Whisper speech recognition writes out every spoken word with a timestamp on each line, ready to export as TXT, SRT or VTT — your first recording is transcribed free, no account, up to 10 minutes. After that, a file costs one minute of balance per audio minute; the original is deleted.
Updated
On this page
How to transcribe audio to text in four steps
To transcribe audio to text here, four steps take a recording from your device to a transcript you can download, with nothing to install and no pausing and rewinding to type it out by hand.
- Choose the file. Drop an MP3, M4A, WAV or another audio file on the upload area, or click it to pick one. Your browser checks the type and size before anything is sent.
- See the price and press Transcribe. A guest's first recording is transcribed straight away and free: in full up to 10 minutes, the first 10 minutes of a longer one. After sign-in the page shows how many minutes the recording will use, and nothing starts until you press Transcribe.
- Check the text against the audio. The recording plays from your own device next to the transcript; click any timestamp and the player jumps to that line.
- Download the transcript. Save it as TXT or Markdown for reading, or as SRT or VTT for subtitles.
No transcription software goes on your computer: the speech to text step runs on our server, and the page only uploads the recording. A long recording shows its progress as minutes done out of the total, so a one-hour interview tells you where it stands instead of leaving a spinner. If you prefer a bigger screen, upload your recording in the workspace, where the transcript, the player and the AI tools sit side by side.
Audio files it accepts, and the limits
This audio to text converter accepts the formats that phones and voice recorders save and editors such as Audacity export, and its only hard limits are size and length.
| Format | Where it usually comes from |
|---|---|
| MP3 | podcast episodes, dictaphones, downloaded lectures |
| M4A (AAC or ALAC inside) | iPhone Voice Memos, Zoom audio-only recordings, QuickTime exports; see how to transcribe an M4A file |
| WAV, AIFF | field recorders, studio software, uncompressed exports |
| FLAC | lossless archives |
| OGG, Opus, AAC | voice messages saved from Telegram, browser recordings |
| MP4, MOV, MKV, WEBM, AVI, MPEG, TS, WMV | video files: the sound track is extracted, the picture is ignored |
Size and length depend on whether you are signed in:
- Without an account: files up to 100 MB; your first recording is transcribed free, in full up to 10 minutes, then short free fragments.
- After sign-in: files up to 500 MB and up to 4 hours long.
- A recording between 100 and 500 MB needs no compressing: sign in, and the same upload goes through.
- A recording longer than 4 hours has to be split into parts first.
Every format takes the same path, so an MP3 to text job and a WAV to text job look and cost the same, and to transcribe audio to text from a video you upload the video itself; the guide to video files and their sound tracks covers MP4, MOV and the size limits.
The extension is only a first check: before you transcribe audio file uploads, the server reads the actual container with ffprobe, so a mislabelled upload is turned away with a plain message before any price is shown.
An audio to text transcription example
A transcript from audio comes back as one line per spoken phrase, and each line opens with the moment that phrase starts, written as minutes and seconds.
Illustrative text in the exact TXT format, not a real recording:
[0:00] Thanks for joining. Today we're going through the first quarter.
[0:05] Sales grew in every region except the north, where two stores closed.
[0:12] Let's start with the online shop, because most of the change came from there.
[0:19] Orders were up by about a fifth compared with the same months last year.Punctuation and capital letters come from Whisper itself, and the spoken language is detected automatically, with no language setting to pick before you start. Past the first hour the stamp grows an hour field, as in [1:02:15]. Stamps mark the start of a phrase rather than each word; the guide to how timestamping a transcription works compares this with four other formats.
What makes a recording transcribe well
Speech recognition is most accurate on one clear voice close to the microphone, and it drops words wherever other sounds compete with speech.
- Helps: a phone or microphone near the speaker, a quiet room, speech without a music bed.
- Hurts: two people talking over each other, background music, traffic, wind, the echo of a large hall.
- Songs: lyrics are transcribed as words, but singing over instruments comes out rougher than speech.
- Distance: a phone on the front desk of a lecture room picks up more than one at the back row.
The same holds for any voice recording to text job: if a passage comes out garbled, the timestamp next to it takes you to the spot, so you can listen again and correct that line by hand.
We publish no accuracy percentage, because it depends on the recording far more than on the software. Your free first recording is the honest test: it shows how your own audio transcribes before you spend anything.
After the transcript: export, subtitles, AI
A finished transcript downloads in four formats at no extra cost, and the AI tools on top of it come at no extra cost too.
- TXT: plain text, with or without the
[m:ss]stamps; a switch on the transcript turns them on and off. - Markdown: the same lines, ready for notes apps such as Obsidian.
- SRT: an SRT file with start and end times, the subtitle file Premiere Pro imports as a caption track.
- VTT: WebVTT captions for HTML5 video players.
Need the subtitles in another format, or the words without the timing? The free in-browser subtitle converters turn SRT into VTT or plain text on your own machine.
AI actions work on the finished transcript: a summary of the recording, detailed notes, an article, a social post, a cleaned-up Pro Transcript, text translation, a chat about the recording or a prompt of your own. They come with the transcribed recording at no extra charge — you pay for minutes of speech recognition, not per button. Each runs only when you press it, and its result can be saved as PDF or DOCX.
What's free and what uses minutes
Your first recording is free without an account — in full up to 10 minutes, the first 10 minutes of a longer one; after that, converting a whole audio file to text costs one minute of balance per minute of audio.
| What | Cost | Conditions |
|---|---|---|
| First recording | free | no account; in full up to 10 minutes, the first 10 minutes of a longer one; on a shared network, a daily limit per network |
| Next recordings without an account | free fragment | a short fragment from the start; the whole recording after sign-in |
| AI analyses | included | every AI analysis of a transcribed recording, no second charge |
| Sign-up gift | 10 minutes | once per account, used first, expire after 7 days |
| Whole file | 1 minute of balance per minute of audio | rounded up to the next minute; the price is shown before you start; charged after a successful transcription |
| Packs | 1 h for $2.99 · 3 h for $7.99 · 10 h for $19.99 · 30 h for $58.99 | one-time payment, no subscription; bought minutes do not expire; refunds within 14 days |
Beyond your first recording, audio to text is free only as a short fragment. The gift covers a recording of up to 10 minutes, such as a short interview transcript or a lecture segment. A 45-minute meeting needs a pack on top of it, and the page tells you so before anything is charged. Past roughly 5 hours of audio every month a subscription can cost less, and the TurboScribe price comparison shows where that line falls.
The text of a YouTube video that already has captions is a different case: it is free for good, with no sign-in and no limits, because it reads captions the video already carries rather than running speech recognition to turn audio into text.
Where your file goes
Audio transcription here keeps only the text: your original file is deleted from our server as soon as the sound has been extracted from it.
- Whenever you transcribe audio to text here, the extracted sound (MP3, 16 kHz, mono) goes to OpenAI’s Whisper for recognition; its pieces are deleted once the transcript is ready.
- Without an account the transcript is not saved on the server: it is held for about an hour until the page collects it, and after that it lives only in your browser.
- An upload that never starts, for example while you top up minutes, is deleted from the server after about an hour.
- After sign-in the transcript is saved in your account.
- The player plays the audio from your device. After a page reload, choose it again to listen; the text is already saved.
- When you pick a file on this page, your browser keeps a temporary copy to hand it to the workspace and deletes it once it has been picked up, or after an hour at most.
The full rules are in the privacy policy.
What it doesn't do
Speaker labels are the main thing this audio to text transcription does not provide: the text is one stream of time-stamped lines, whoever is speaking.
- No stamps per word and no stamps at fixed intervals: one per phrase.
- No subtitles burned into a video; you get SRT or VTT files to load in an editor.
- No recording from a microphone and no live voice to text dictation; upload a finished recording.
- No links to Google Drive, Dropbox or Spotify; download it first, then upload it here.
- Files over 500 MB or longer than 4 hours are turned away; split them first.
- Only the first audio track is used. A recording with separate tracks per person needs a mixed track, or each track uploaded on its own.
What you can count on instead when you transcribe audio to text here: the price is shown before the start, minutes are charged only after a successful transcription, and the export of a finished transcript is included.
Guides by format and task: M4A to text · Timestamping transcription
Frequently asked questions
Can I transcribe audio to text for free?
Yes, your first recording: without an account it is transcribed free, in full if it runs up to 10 minutes and the first 10 minutes if it is longer. After that a guest gets a short free fragment; the whole recording uses minutes from your balance, and signing up adds 10 of them.
Do I need an account?
No for your first recording and for short free fragments. Yes for whole files after that: sign-in works through a link sent to your email, with no password to create.
Which audio formats and sizes can I upload?
This audio to text converter takes MP3, M4A, WAV, FLAC, OGG, Opus, AAC and AIFF, plus video files such as MP4, MOV and MKV. Up to 100 MB without an account; up to 500 MB and 4 hours after sign-in.
Is my audio file stored?
No. The original is deleted from the server right after the sound is extracted, or after about an hour if the transcription never starts, and the audio pieces are deleted after recognition. The text stays in your browser, or in your account if you are signed in.
Can it tell speakers apart?
No. There are no speaker labels: the transcript is one stream of time-stamped lines, whoever is talking.
What happens when free transcription is unavailable?
On a shared network, such as an office or a mobile carrier, free transcription has a daily limit per network, and the page says so when it is reached. You then get a short fragment, and the whole recording after sign-in, with the 10 gift minutes.