MediaScribe

Timestamps

Timestamping transcription: formats, a sample and automatic timestamps

Timestamping a transcription puts the moment each phrase was spoken in front of it — [12:04] like that — and here every line of an uploaded recording gets its start time automatically, clickable in the player and kept in TXT, SRT and VTT exports; your first recording is free without an account, in full up to 10 minutes. Five formats, a time coded transcript sample and the limits of automatic stamps follow.

Updated

Without an account: files up to 100 MB; the first transcription is free in full up to 10 min, then a short free fragment. After sign-in: up to 500 MB and 4 hours. · 1 min of balance per minute of recording
MP3 · M4A · WAV · FLAC · OGG · MP4 · MOV · MKV · WEBM · AVI · SRT · VTT
On this page

What timestamps in a transcription are for

Timestamps in a transcription tie each line to a moment in the recording, so a reader goes from a quote straight back to the sound.

  • Checking a word: a doubtful name or number is verified by jumping to its time instead of replaying the whole file.
  • Quoting with a source: journalists and researchers cite “interview 3, 14:52” so anyone reading the interview transcript can find the passage.
  • Subtitles: every caption in a subtitle file needs a start and an end time.
  • Coding interviews: in qualitative research transcription, tools such as NVivo and MAXQDA link a time coded transcript to its audio.

Style guides for time coded transcription use several kinds of timestamps:

  • per phrase or sentence;
  • at a fixed interval, for example every 30 or 60 seconds;
  • at every change of speaker;
  • next to unclear passages, as in [inaudible 00:12:31].

Time stamping in transcription by hand means pausing the audio at every stamp, reading the player’s clock and typing it in, hundreds of times for an hour of speech. Automatic timestamping transcription on this page is the first kind: one stamp per phrase.

Timestamp formats, side by side

The timestamping format in transcription depends on where the text goes next, and five forms cover almost every case.

FormatLooks likeExpected byHere
minutes:seconds[1:15], [1:02:15] past an hourreaders, notes, show notesyes (TXT)
hh:mm:ss with leading zeros[00:01:15]many transcription style guidesno
Markdown line- **[1:15]** textnotes apps, docsyes (MD)
SRT00:01:15,000 --> 00:01:19,200Premiere Pro, YouTube caption uploadsyes
WebVTT00:01:15.000 --> 00:01:19.200HTML5 video playersyes

SRT to VTT is a small change: the two differ mainly in the comma or dot before the milliseconds and in the WEBVTT header, so it takes one click to convert SRT timing to WebVTT when a player asks for it. A time coded transcript meant for reading uses [1:15]; one meant for a video player uses SRT or WebVTT.

A time coded transcript example

A time coded transcript in plain text puts each phrase on its own line after its start time, and the subtitle version of the same opening adds an end time.

Illustrative example, invented text rather than a real recording, in the exact format of the TXT export:

[0:00] Okay, it’s recording. Could you start by telling me what you do?
[0:04] Sure. I run the repair workshop on the ground floor, mostly bikes and scooters.
[0:09] How long have you been in this building?
[0:14] Six years in March. Before that we had a garage across the river.
[0:21] And what changed when you moved here?
[0:26] More walk-in customers, for a start. People see the window and just come in.
[0:33] Is that most of your work now?
[0:38] About half. The rest is delivery companies bringing their fleets in.

The first three lines of the same example as an SRT export:

1
00:00:00,000 --> 00:00:04,200
Okay, it’s recording. Could you start by telling me what you do?

2
00:00:04,200 --> 00:00:09,800
Sure. I run the repair workshop on the ground floor, mostly bikes and scooters.

3
00:00:09,800 --> 00:00:14,700
How long have you been in this building?

This timecode transcription example uses minutes and seconds, so [0:14] means 14 seconds in. The SRT version adds an index and an end time to each line, which a video player needs to hide the caption again.

How to add timestamps automatically

Timestamping transcription automatically takes an upload and nothing else: the recording goes in, and the lines come back already stamped.

  1. Upload the recording. Drop an audio or video file on the upload area at the top of the page.
  2. Let Whisper split the speech. Whisper divides the speech into phrases and records the moment each one starts.
  3. Check a line in the player. Click a timestamp and the recording jumps to that moment, so a doubtful word is checked in seconds.
  4. Choose timestamps on or off. A switch shows or hides the stamps for TXT and Markdown; SRT and VTT always keep them.
  5. Export. Download TXT, Markdown, SRT or VTT.

It works the same for an interview in MP3, a lecture filmed as MP4 or an M4A recording from a phone. For the formats, size limits and prices in one place, see how to upload any audio file for a transcript.

Where automatic timestamps are exact, and where they aren't

Automatic timestamping transcription is reliable for finding a phrase and approximate at the edges of it.

  • Per phrase, not per word. Each transcription timestamp marks where a phrase begins; a word in the middle of a long sentence has no time of its own.
  • Phrase edges are approximate. Whisper’s speech recognition decides where one phrase ends and the next starts, and a pause or a fast speaker shifts that boundary.
  • No interval stamps and no speaker-change stamps. A research template that asks for a time every 60 seconds or at each turn has to be adjusted by hand.
  • One fixed transcript format. The text export writes [1:15], not [00:01:15].

For subtitles and for “find the moment someone said this”, a transcript with timestamps per phrase is enough. For verbatim transcription of legal records with a stamp every minute, it is a starting draft.

Timestamps to chapters, subtitles and quotes

A time coded transcript can turn into subtitles, chapters or clean quotes, depending on which export you take.

  • Subtitles: the SRT or VTT export loads straight into an editor or player.
  • Clean text: switch the stamps off before export, or strip the timecodes from an SRT file you already have.
  • Chapters: the stamps show where each topic starts, so chapter markers for a video description or a podcast page can be written straight from the transcript.
  • Quotes: copy a line together with its stamp, and the reader can find the passage in the recording.

If the recording is a YouTube video with captions, there is nothing to upload: the timestamps for a YouTube video with captions come free, straight from the caption track.

What a timestamped transcript costs

Timestamping transcription costs nothing extra: the stamps come with every transcript, and the transcript itself is priced by length. There is no separate option to buy or switch on.

Your first recording is free without an account, in full up to 10 minutes. Signing up adds 10 minutes once, and after that a file uses one minute of balance per minute of audio, with the price shown before the start. The full breakdown of packs is on the audio to text transcription page.

Frequently asked questions

How do you timestamp a transcription?

By hand, you pause the audio and type the current time at each line or interval. Automatically, you upload the recording and get a transcription with timestamps: every line carries its start time.

How often should a transcript have timestamps?

It depends on who reads it: subtitles need a time on every line, while research transcripts often ask for one at each speaker turn or at a fixed interval. The transcripts here have one per phrase.

Are the timestamps word-by-word?

No. There is one timestamp per phrase, marking where that phrase starts.

Can I remove the timestamps?

Yes. Switch them off before exporting TXT or Markdown. SRT and VTT always keep them, because players need the timing.

Can it stamp every speaker change?

No. The transcript has no speaker labels, so it cannot mark who starts talking when.

Is it free?

Yes for your first recording: without an account it is free in full up to 10 minutes, and the first 10 minutes of a longer one. After that a file uses minutes from your balance, and signing up gives 10 minutes.

Do I need an account?

No for your first recording: timestamping transcription works without an account for files up to 100 MB. For the whole recording, yes: sign-in works through a link sent to your email, adds 10 gift minutes that expire after 7 days, and raises the limit to 500 MB and 4 hours.

Is my recording stored?

No. Your browser keeps a temporary copy for at most an hour to hand the file to the workspace. On the server the original is deleted once the sound is extracted, or after about an hour if the transcription never starts. The transcript is saved in your account only when you are signed in; without an account it stays in your browser.

Get the transcript now

Upload a recording at the top of the page: every line comes back with its start time, and your first recording is free without an account, in full up to 10 minutes.