Convert · VTT → TXT
VTT to TXT converter
Zoom and Teams hand you a caption track rather than a readable page; converting VTT to TXT closes that gap, because the WEBVTT header, the cue settings and every timecode come out, and the conversation is left as continuous prose. Four sections follow: what is discarded, where these tracks come from, what the wording is good for, and what the converter cannot do.
Updated
Convert a VTT file to TXT in one pass
Drag the WebVTT track onto the converter above, or paste its contents into the left pane, and readable prose appears on the right. No upload, no account, no ceiling on how many you run: the page does the work in your browser and forgets it when the tab closes. A 45-minute meeting usually runs to 700 or 800 cues, and all of them resolve before you let go of the mouse. Copy the output, or download a .txt for your notes app.
Holding SubRip rather than WebVTT? The SubRip version of this strip behaves identically.
What is discarded
WebVTT wraps the dialogue in scaffolding, and every part of it is scaffolding nobody reads:
- the opening
WEBVTTline and any header metadata after it; - cue identifiers — the optional label sitting above a timing line;
- timing rows such as
00:07:02.400 --> 00:07:05.960, with settings likeline:,position:andalign:; NOTEcomments andSTYLEblocks, written for the player and never for the reader.
What remains is the wording, in its original order. Prefer to keep the cue boundaries and change only the format? Hold on to the timings and switch to SubRip instead.
Where these tracks come from
Most people meet WebVTT without asking for it. Zoom writes one for every transcribed cloud recording; Microsoft Teams offers the same download beside a meeting; automatic captioning services and streaming platforms hand you a .vtt because it is the format their delivery pipeline already speaks. Lecture-capture systems and podcast hosts add to the pile, and none of them offers a readable version next to the machine-made one. So people arrive here holding a caption track when what they wanted was the conversation — a stand-up, a webinar, a customer call. Stripping the scaffolding turns a delivery artefact back into something a colleague who missed the call can read in one sitting.
What the wording is good for
Plain text goes where captions cannot: into a search index, a quotation in a report, a paragraph of minutes, a glossary of terms used in a webinar, or an assistant asked for a summary. It is the cheapest possible input for anything automated, because exact words need no decoding, and word counts, reading time and term frequency all become computable the moment the timecodes go. One caveat is worth budgeting for — machine captions arrive without reliable punctuation and often without capital letters, so allow a minute of editing before the text faces an audience. Working from a link rather than a download, you can turn a YouTube video into text without a file at all.
The limits, stated plainly
The converter removes markup and does nothing else: no translation of the track you paste, no media uploads, no speaker diarisation, no punctuation repair. Cues are joined in the order they appear, so an overlapping pair of live captions still reads as an overlap, and two people talking at once stay tangled on the page exactly as they were in the room. Everything is one-way — save the original before stripping it if a timed version might be wanted again. Starting from a script instead of a recording? Build a WebVTT skeleton out of plain lines from the opposite direction.
Frequently asked questions
How do I convert VTT to TXT?
Drag the .vtt onto the converter above or paste its text. The WEBVTT header and every timing line are removed, leaving clean prose on the right to copy or save as a .txt. Free, and no sign-in.
Does it remove the timestamps?
Yes, that is the job: cue boundaries and positioning go, wording stays. Keep the original file if a timed track might be needed again.
How do I read a Zoom transcript as plain text?
A cloud recording gives you a .vtt. Paste it in and the conversation comes back as continuous text, with whatever speaker labels the recording wrote into the cue wording still in place.
Is anything uploaded?
No. The page reads the track locally, so a confidential meeting transcript never leaves your computer.