Video to text
Transcribe video to text: upload the file, get a timestamped transcript
To transcribe video to text, only the sound track matters: drop an MP4, MOV, MKV or WEBM file below, and the speech is taken from the video, recognised by Whisper and written out line by line with timestamps, ready for TXT, SRT or VTT — your first recording is free without an account, up to 10 minutes. After that, a minute of balance covers a minute of recording.
Updated
On this page
What video transcription reads: the sound track, not the picture
Video transcription is speech recognition applied to a video’s sound track: ffmpeg, the open-source media toolkit, copies the first audio track out of the file without decoding the picture, reduces it to a 16 kHz mono MP3 and hands it to OpenAI’s Whisper model, so the image codec, H.264 or HEVC, makes no difference.
- The picture is ignored. Slides, captions drawn into the frame and other on-screen text are not read; only what is said ends up in the transcript.
- No sound, no text. A screen recording made without a microphone or system audio is turned away with the message “There is no sound in this file.”
- One track. When a file carries several audio tracks, such as a film with two dubs, the first one is used.
- Too big? Upload the sound alone: QuickTime Player can save a clip as audio only, and the exported file is transcribed like any M4A voice memo. The text comes out the same, and the file shrinks to a fraction.
Because the path is identical, a clip and an audio file cost the same per minute, and transcribing audio to text from a recording without a picture takes exactly this route.
How to transcribe a video to text in four steps
Four steps take a file from your computer or phone to a transcript you can download, and that is the whole way to transcribe video to text here: nothing to install, no rewinding to type out what was said.
- Choose the video. Drop an MP4, MOV, MKV, WEBM or another file on the upload area, or click it to pick one. Your browser checks the type and size before anything is sent.
- Start the transcription. A guest's first recording starts straight away and is free: in full up to 10 minutes, the first 10 minutes of a longer one. After sign-in the page shows how many minutes the file will use, and nothing runs until you press Transcribe.
- Watch and read side by side. The clip plays from your own device next to the transcript; click any timestamp and the player jumps to that moment.
- Download the text or the subtitles. Save TXT or Markdown to read and edit, SRT or VTT to use as subtitles.
The speech recognition runs on our server, so the page needs no plugin and no desktop app. A long recording reports progress as minutes done out of the total: a 90-minute webinar shows where it stands instead of a spinner. For a larger screen, open the workspace with your file, where the player, the transcript and the AI tools sit side by side.
Video files it accepts, and the size limits
This video to text converter accepts the containers that phones, cameras, meeting apps and screen recorders save, and its only hard limits are file size and length.
| Format | Where it usually comes from |
|---|---|
| MP4 | Android cameras, Zoom local recordings, downloaded Teams and Google Meet recordings, Loom; see MP4 to transcript for phone and meeting recordings |
| MOV | iPhone camera, screen recordings on a Mac, QuickTime Player |
| MKV, WEBM | OBS Studio, browser-based screen recorders |
| AVI, WMV, MPEG, TS | older cameras and Windows software, TV and broadcast captures |
Size and length depend on whether you are signed in:
- Without an account: files up to 100 MB; your first recording is transcribed free, in full up to 10 minutes, then short free fragments.
- After sign-in: files up to 500 MB and up to 4 hours long.
- A file longer than 4 hours has to be split into parts first.
Picture weighs far more than speech: an hour of 1080p screen recording can exceed the 500 MB cap, while the same hour of sound takes tens of megabytes. The file extension is only a first check; the server reads the real container with ffprobe and turns away anything that is not a media file before a price is shown.
Which videos transcribe well
Talking-head videos, lectures, interviews, webinars and screen recordings with narration transcribe best, because one clear voice sits close to the microphone.
- Works well: a clip-on or headset microphone, a quiet room, a single speaker at a time.
- Works worse: background music under speech, crowd noise, people talking over each other, wind on an outdoor phone clip.
- Music videos: sung lyrics come out rougher than speech.
OpenAI trained Whisper on 680,000 hours of audio collected from the web, as its 2022 paper “Robust Speech Recognition via Large-Scale Weak Supervision” describes, so accents and casual speech are within what it was built for; overlapping voices and music under speech remain the usual causes of errors. We publish no accuracy percentage: it depends on the recording far more than on the software. Your free first recording is the honest test of video transcription on your own footage, and the timestamp next to a garbled line takes you straight to the spot to listen again.
A video transcript example
A video transcript comes back as one line per spoken phrase, and each line opens with the minute and second at which that phrase starts.
Illustrative text in the exact TXT format, not a real recording:
[0:00] Hi everyone, this is the product demo for the new dashboard.
[0:06] On the left you can see the filters we added last week.
[0:13] Let me open the sales report, because that's where most questions came from.
[0:21] As you can see, the totals now update without reloading the page.Punctuation and capitals come from Whisper, and the spoken language is detected on its own, with nothing to set before the start. After the first hour the stamp gains an hour field, as in [1:04:10]. Each stamp marks the start of a phrase, not of every word; the guide to timestamping transcription formats compares this with four others.
After the transcript: subtitles, export, AI
A finished video transcript downloads in four formats at no extra cost, and two of them are subtitle files.
- TXT and Markdown for reading, notes and editing, with or without the
[m:ss]stamps. - SRT for Premiere Pro, DaVinci Resolve and YouTube Studio; VTT for HTML5 players.
The same upload gives video to SRT subtitles with start and end times from the recognition, so no second pass is needed. To re-time, convert or strip the result, the subtitle converters that run in your browser handle SRT, VTT and plain text.
AI actions work on the finished transcript: a summary, detailed notes, key moments, chapters with timecodes, an article, a social post, a cleaned-up Pro Transcript, a chat about the recording or a prompt of your own. They come with the transcribed recording at no extra charge: you pay for minutes of speech recognition, not per button.
Free video transcription, and what uses minutes
Free video transcription here covers your first recording without an account, in full up to 10 minutes and the first 10 minutes of a longer one; after that, transcribing a whole file costs one minute of balance per minute of recording, from $0.033 per minute.
| What | Cost | Conditions |
|---|---|---|
| First recording | free | no account; in full up to 10 minutes, the first 10 minutes of a longer one; on a shared network, a daily limit per network |
| Next videos without an account | free fragment | a short fragment from the start; the whole file after sign-in |
| Export and AI analyses | included | TXT, MD, SRT, VTT and every AI analysis of a transcribed file, no second charge |
| Sign-up gift | 10 minutes | once per account, used first, expire after 7 days |
| Whole file | 1 minute of balance per minute of recording | rounded up; the price is shown before you start; charged after a successful transcription |
| Packs | 1 h for $2.99 · 3 h for $7.99 · 10 h for $19.99 · 30 h for $58.99 | one-time payment, no subscription; bought minutes do not expire; refunds within 14 days |
In practice the gift covers a 10-minute screen recording or a short interview. A 75-minute lecture needs a pack on top of it, and the page says so before anything is charged. The “first recording” is one per guest: a YouTube video without captions transcribed earlier counts as that recording too.
Does a video already on YouTube need uploading?
No: most YouTube videos carry captions, and their text is read directly, with no upload and no speech recognition.
The free text of a YouTube video with captions comes from the caption track, with no sign-in and no limits, and an AI summary of such a video is free as well. Speech recognition, and the minutes it uses, is needed only when there is no text to start from: a file of your own, or a YouTube video without captions.
Where your video goes
Only the text is kept when you transcribe video to text here: your video file itself is deleted from our server as soon as its sound has been extracted.
- The extracted sound goes to OpenAI’s Whisper for recognition; its pieces are deleted once the transcript is ready.
- Without an account the transcript is not saved on the server: it is held for about an hour until the page collects it, then lives only in your browser.
- An upload that never starts, for example while you top up minutes, is deleted after about an hour.
- After sign-in the transcript is saved in your account.
- The player plays the file from your own device. After a page reload, choose the file again to watch it; the text is already saved.
- When you pick a file on this page, your browser keeps a temporary copy to hand it to the workspace and deletes it once it has been picked up, or after an hour at most.
The full rules are in the privacy policy.
What it doesn't do
On-screen text is the main thing this video to text transcription leaves out: it turns speech into text and does not read the picture.
- No speaker labels; one stream of time-stamped lines.
- No subtitles burned into the picture; you get SRT or VTT files to load in an editor or player.
- No extraction of subtitle tracks already stored inside an MKV or MP4; the text is generated from speech.
- No links to Google Drive, Dropbox, Vimeo or TikTok; download the file first, then upload it here.
- Files over 500 MB or longer than 4 hours are turned away; upload the sound alone or split the file.
- Only the first audio track is used.
What you can count on instead: the price is shown before the start, minutes are charged only after a successful transcription, and the export of a finished transcript is included.
Guides by format and task: MP4 to transcript · Video to SRT
Frequently asked questions
Can I transcribe a video to text for free?
Yes, your first recording: without an account it is transcribed free, in full if it runs up to 10 minutes and the first 10 minutes if it is longer. After that a guest gets a short free fragment; the whole file uses minutes from your balance, and signing up adds 10 of them.
Do I need an account?
No for your first recording and for short free fragments. Yes for whole files after that: sign-in works through a link sent to your email, with no password to create.
Which video formats and sizes can I upload?
MP4, MOV, MKV, WEBM, AVI, MPEG, TS and WMV, plus audio files such as MP3, M4A and WAV. Up to 100 MB without an account; up to 500 MB and 4 hours after sign-in.
My video is bigger than the limit. What can I do?
Upload the sound on its own. Export the audio track, for example as an M4A from QuickTime Player, and upload that file: the transcript is the same, because only the sound is ever recognised, and the file is a fraction of the size.
Is my video stored on your server?
No. The file is deleted from the server right after its sound is extracted, or after about an hour if the transcription never starts, and the audio pieces are deleted after recognition. The text stays in your browser, or in your account if you are signed in.
Does it read text shown on screen or burned-in subtitles?
No. Only speech is turned into text. Slides, captions drawn into the picture and any other on-screen text are not read.
Can it tell speakers apart?
No. There are no speaker labels: the transcript is one stream of time-stamped lines, whoever is talking.