Reference · WebVTT
What is a VTT file?
Every VTT file announces itself on line one: the word WEBVTT, followed by timed cues that a browser paints over an HTML5 video — plain text, nothing more. What the rest of the document holds, which services produce these tracks, how to open and validate one, how it attaches to a page, and where SubRip still wins are covered in turn.
Updated
What is inside a VTT file
The first line is the giveaway: every WebVTT document begins with the word WEBVTT, and a player that does not find it discards the track entirely. Cues follow, separated by blank lines. A cue may open with an optional identifier, then carries a timing row such as 00:01:30.500 --> 00:01:33.000 — a dot before the milliseconds, with the hour field allowed to go missing — followed by the text to display.
The format has room for more than dialogue. Cue settings appended to the timing row (line:, position:, align:, size:, vertical:) place the words on screen; NOTE blocks hold comments; STYLE blocks carry CSS aimed at the ::cue pseudo-element; REGION definitions describe the roll-up areas live captioning uses. Encoding is not negotiable, because the specification requires UTF-8. If the conversation is what you were after, reduce the cues to plain prose and read it like any other document.
Where these tracks come from
Most people meet the format without asking for it. Zoom writes one alongside every transcribed cloud recording; Microsoft Teams offers the same download after a meeting; automatic captioning services export it by default. Streaming is the other large source, since an HLS playlist delivers subtitles as WebVTT segments, so anything played adaptively in a browser is already using this format underneath. Video platforms and course tools follow suit, because it is what their players read without conversion — which is why a .vtt tends to arrive unrequested rather than chosen. Zoom and Teams both write speaker names into the cue text itself, so the strip carries them over into the finished prose.
Opening and validating one
It is plain text, so Notepad, TextEdit or VS Code opens it in an instant, and VLC attaches it to a playing video when you drag it in. Subtitle Edit and Aegisub add a waveform when cues need nudging. To check that a track is genuinely valid, the quickest test is the browser itself: point a <track> element at it, open the developer console and watch for parse warnings.
Three faults account for nearly every rejected track — a missing or misspelled header, commas where the dots should be, and a file saved in a legacy code page instead of UTF-8. All three are visible on the first screenful. A track written for an editing timeline rather than a browser is better off as SubRip for an editing timeline.
Using it on a web page
Attaching captions takes one element inside your video: <track kind="subtitles" src="talk.vtt" srclang="en" label="English" default>. The kind attribute decides what the track is — subtitles and captions for dialogue, chapters for a navigable outline, metadata for data your own script reads. Two server-side details cause most silent failures: the response must arrive with the text/vtt MIME type, and a track hosted on another domain needs CORS headers, or the browser drops it without a word in the interface. Already holding a SubRip copy? Make a WebVTT track from an SRT you already have and the timings carry over unchanged.
WebVTT next to SubRip
They overlap almost entirely: both are plain-text lists of timed cues, and a converter moves between them in milliseconds. WebVTT adds the header, the dot, mandatory UTF-8, optional cue identifiers, positioning and styling. SubRip keeps compulsory numbering, a comma, and the widest support of any subtitle format in existence. Use WebVTT when a browser or a stream will play the video, SubRip when an editor, a desktop player or an upload form is involved. When the source is a published video and neither format is really wanted, YouTube captions as plain text gets you the wording on its own.
Frequently asked questions
What is a VTT file?
A caption track in the WebVTT format, defined by the W3C for HTML5 video. It is plain text: a WEBVTT header, then cues that each carry a start time, an end time and the words to display.
How do I open a VTT file?
Any text editor opens it, since it is plain text. Subtitle Edit and Aegisub add a waveform for re-timing, and dragging it onto a VLC window attaches it to whatever is playing.
What is the difference between a VTT file and an SRT file?
WebVTT requires its header, uses a dot before milliseconds, must be UTF-8 and supports positioning, regions and CSS styling. SubRip numbers every cue, uses a comma, and is the format editors and desktop players expect.
Where do VTT files come from?
Zoom and Teams recordings, automatic captioning services, streaming platforms and HLS playlists all produce them, because WebVTT is the format browsers and delivery pipelines speak natively.