MediaScribe
🇬🇧English
Add to Chrome

Convert · TXT → VTT

TXT to VTT converter

Captions bound for a web page have to be WebVTT, and converting TXT to VTT is how a script gets there: each line becomes a cue of the length you choose under a valid WEBVTT header, ready to re-time against the audio. From there the page explains the timing catch, the shape of the output, attaching it to a video element, and the limits.

Updated

TXT→VTTruns in your browser · nothing uploaded

Input — TXT

⏱ No timings in plain text — each line becomes a cue of . A subtitle skeleton you re-sync in any editor; we don’t guess real timings.

Output — VTT

Convert a TXT file to VTT in one pass

Paste your script into the converter above — one caption per line — or drag a .txt onto it, set how long each line should hold, and download the .vtt. The page works locally and keeps nothing: a draft narration for an unannounced product stays on your machine, and you can regenerate the file as often as the script changes.

New to the format? What a valid WebVTT document must contain explains why the first line matters so much.

The timing catch, said out loud

A plain script carries no timing data at all, so nothing here can work out when a line belongs on screen — and it will not invent an answer. Every line is given the fixed duration you chose, running consecutively from zero. The download is a skeleton: correct structure, your exact wording, placeholder timings waiting to be moved.

It is a head start, not a finished caption track. Open it in an editor and align the cues with the audio — the single part of captioning that needs a human listening, and the part no text-to-caption converter can honestly claim.

What the skeleton looks like

The output opens with WEBVTT on its own line, then carries one cue per line of your script: a timing row such as 00:00:04.000 --> 00:00:06.000 followed by the words. Milliseconds use a dot, as the specification requires, and the encoding is UTF-8, so accented characters and non-Latin scripts survive the round trip. Cue identifiers, positioning settings and STYLE blocks are left out deliberately, because they are optional and a skeleton is easier to re-time without them. The result is ordinary text, so two drafts of a narration can be diffed in Git, or a colleague can fix a typo in a browser tab, long before anybody opens a timeline. Want the editor-friendly twin? The SubRip version of this skeleton numbers each cue and uses commas.

Putting the track on a page

Once the cues are synced the track attaches with one element: <track kind="subtitles" src="intro.vtt" srclang="en" label="English"> inside your <video>. Serve it from the same origin, or add CORS headers if it lives on a CDN, and make sure the server sends text/vtt — a wrong MIME type is the usual reason a perfectly good track never appears. The same document can then be reused as a chapter list or a metadata track, since WebVTT covers those as well. If the words really come from a published video, save the existing subtitles instead of retyping them.

The limits, stated plainly

There is no listening, no translation and no media handling in this tool: text goes in, a caption skeleton comes out. Long lines are not re-wrapped for you, overlapping speech is not detected, and durations are uniform by design rather than by accident. Everything it does promise is verifiable in a text editor within seconds. For a finished, already-timed transcript of something published online, take the finished wording from the video itself and skip the skeleton entirely.

Frequently asked questions

How do I convert TXT to VTT?

Paste your script into the converter above with one caption per line, pick how many seconds each should hold, and download the .vtt. Free, no sign-in, and nothing leaves your browser.

Can it detect the real timings?

No. A script has no clock in it, so every cue gets the fixed duration you selected. Aligning those cues with speech happens afterwards in an editor, by ear.

Why choose VTT instead of SRT?

Because the captions are headed for a web page. The HTML5 track element reads WebVTT only, and streaming players expect it too. For a timeline in Premiere or Resolve, SubRip remains the safer pick.

Does the output start with the WEBVTT line?

Yes. The header is written for you, timecodes use dots, and the text is UTF-8, which is what the specification demands, so the document validates as it stands.