SRT
Video to SRT: subtitles generated from the speech in your file
Video to SRT here means subtitles written from the speech, not typed by you: upload an MP4, MOV or MP3, and each spoken phrase becomes a numbered cue with its start and end time, ready for Premiere Pro, DaVinci Resolve or YouTube Studio — your first recording is free without an account, up to 10 minutes.
Updated
On this page
How to turn a video into an SRT file
Four steps take a recording to finished subtitles, and none of them involves typing a timecode.
- Upload the video or audio. Drop an MP4, MOV, MKV, MP3, WAV or another media file on the upload area at the top of the page.
- Let the speech be recognised. A guest's first recording starts straight away and is free, in full up to 10 minutes. After sign-in the page shows the price in minutes, and nothing runs until you press Transcribe.
- Check the lines against the player. Read the transcript next to the player and click a timestamp to hear any line again.
- Export SRT or VTT. Choose SRT for an editor or YouTube Studio, VTT for a web player. Both carry start and end times.
SubRip is one of four exports of the same transcript. If you mainly need the text of the video, start from full video transcription instead; the captions come with it.
What the generated SRT looks like
Generated subtitles are plain text: a cue number, a timing line with start and end separated by an arrow, and the words of one phrase.
Illustrative cues in the exact SubRip format, not a real recording:
1
00:00:00,000 --> 00:00:05,400
Hi everyone, this is the product demo for the new dashboard.
2
00:00:05,400 --> 00:00:12,100
On the left you can see the filters we added last week.Milliseconds follow a comma, as SubRip requires; the VTT export writes them after a dot and opens with a WEBVTT line. For the rules of the format itself, see what an SRT file is.
How the cues are cut, and what to fix by hand
Each cue is one phrase exactly as Whisper speech recognition split the audio, timed from where the phrase starts to where it ends.
- Short phrases make comfortable subtitles; a long uninterrupted sentence becomes one long cue.
- There is no setting for characters per line, number of lines or reading speed.
- Phrase edges are approximate: a pause or a fast speaker shifts them by a fraction of a second.
Streaming services set such limits in their style guides: Netflix’s English timed text guide allows 42 characters per line and two lines per cue. Captions held to a standard like that are best finished in a subtitle editor such as Subtitle Edit or Aegisub: the timings are already there, so the work is splitting and trimming, not timing from zero.
Loading the subtitles into an editor or a player
Cues generated from speech open in every tool that accepts SubRip, because it is the same format a person would type.
- Premiere Pro: import the cues and drag them onto the timeline as a caption track.
- DaVinci Resolve: File → Import → Subtitle.
- YouTube Studio: Subtitles → upload with timing.
- VLC: give the subtitle file the same name as the video and keep both in one folder; it loads automatically.
- A web page: use the VTT export in a
<track>element, or convert the SRT to WebVTT later.
Three ways to get an SRT, and which one is yours
Which route fits depends on what you start from: captions on YouTube, a typed script, or only a recording.
| You have | Route | Cost |
|---|---|---|
| A YouTube video with captions | Export the caption track | free, no sign-in |
| A script, but no timings | Text to SRT skeleton, then sync by hand | free, in your browser |
| A recording and nothing else | Speech recognition writes the text and timings (this page) | first recording free, then minutes |
An SRT from YouTube subtitles covers the first case for free, and the TXT to SRT converter builds a skeleton from a typed script for the second. Only the third needs speech recognition.
Audio to SRT works the same way
Audio to SRT follows the identical path, because a clip is reduced to its sound before recognition anyway.
A podcast episode, a voice-over or a lecture recorded on a phone gives subtitles just as an MP4 does: MP3 to SRT, M4A to SRT and WAV to SRT all end in the same export. Upload the recording through audio file transcription and pick the subtitle export at the end; MP4 recordings are covered on the page about MP4 transcripts.
What an SRT from speech costs
The export itself is free; what costs minutes is the speech recognition behind it, one minute of balance per minute of recording.
Your first recording is free without an account, in full up to 10 minutes. A 10-minute tutorial then costs 10 minutes, which the 10-minute sign-up gift covers, and the price always shows before the start.
What it doesn't do
Video to SRT here generates subtitles from speech; it does not edit, burn or extract them.
- No subtitles burned into the video picture.
- No extraction of subtitle tracks already inside an MKV or MP4.
- No speaker names and no sound descriptions such as [music].
- No line-length or reading-speed settings.
- Files over 500 MB or 4 hours are turned away.
Frequently asked questions
Can I convert video to SRT for free?
Yes, your first recording: without an account it is transcribed free, in full up to 10 minutes and the first 10 minutes of a longer one, and the subtitle export is included. After that a file uses minutes from your balance; signing up adds 10.
Do I need an account?
No for your first recording, with files up to 100 MB. Yes for whole files after that: sign-in by an email link raises the limit to 500 MB and 4 hours.
Can I make an SRT from an MP3 or other audio file?
Yes. Audio to SRT works the same way as video: upload the MP3, M4A or WAV and export the subtitles.
Will the subtitles be burned into my video?
No. You get a separate subtitle file; your editor or player shows it over the video.
Is my file kept on your server?
No. The original is deleted right after its sound is extracted, or after about an hour if the transcription never starts. The subtitles stay in your browser or in your account.
Can I set the number of characters per line?
No. Each cue is one phrase as the recognition split it; long phrases are best shortened in a subtitle editor.