Guide

How to Write a Transcript of a Video (3 Methods, One Free)

Transcribing a video by hand takes about four times its runtime — roughly 40 minutes of typing for a 10-minute video. There are three faster routes, and which one you want depends on whether you need the words, or the words plus their timings. That second distinction matters more than most guides admit.

First decide: transcript or subtitles?

A transcript is the plain text of what was said. Subtitles are that same text split into short lines, each carrying a start and end timestamp — normally an .srt or .vtt file.

Every subtitle file contains a transcript. A transcript does not contain timings, and you can't recover them later without redoing the work. So unless you're certain you only need the words, export .srt. It costs nothing extra at the time and saves the whole job later.

Method 1 — By hand

Play, pause, type, rewind. Accurate, and appropriate when the audio is poor, the terminology is unusual, or the transcript is going into something where an error is expensive.

Two things make it bearable: slow the playback to 0.75×, and use a player with keyboard-only seek so your hands never leave the keyboard. Expect ~4× runtime.

Method 2 — YouTube auto-captions (free)

If the video is on YouTube — yours or someone else's — the transcript already exists:

  1. Open the video on desktop.
  2. Click the three dots below the player.
  3. Choose Show transcript.
  4. The panel opens beside the video. Toggle timestamps off if you just want the prose, then select and copy.

For your own uploads you can go further: in YouTube Studio, open Subtitles, and download the track as .srt — that gives you the timings too. The YouTube-specific routes (desktop, mobile, Studio, and what to do when the button is missing) are covered step-by-step in our guide on how to get a YouTube transcript.

The caveat is accuracy. Auto-captions are speech recognition, and they fail in a predictable pattern: proper nouns, numbers, and jargon. Those are also the words that carry the most meaning, so never publish auto-captions unread.

Method 3 — Speech-to-text tools

Any dedicated transcription tool will take an audio or video file and return text, usually with timings, usually in a couple of minutes. What separates them in practice is not headline accuracy but three things:

If you're transcribing your own narration, you likely already have the audio isolated, which is the best case for recognition accuracy — no music bed, one speaker, consistent level.

Cleaning up the result

Work in this order. It's fastest and it front-loads the errors that matter.

  1. Proper nouns and numbers first. Highest error rate, highest cost when wrong.
  2. Punctuation and paragraph breaks. Recognition output tends to arrive as one long run.
  3. Filler words. Remove for reading transcripts; keep for captions, where they match the audio the viewer hears.
  4. Line lengths, if it's subtitles. Aim for one short clause per line — roughly 40 characters — so nothing gets cropped on a phone.

What an .srt file is actually good for

Most people treat subtitles as an accessibility chore and stop there. But an .srt is a structured record of what was said and exactly when — which is useful well beyond captions:

That last point is worth spelling out. The hardest part of editing a narrated video isn't the words — it's deciding what should be on screen for each line, then producing it. But an .srt already contains every line and its exact timing. That is the shot list. Ahacut reads it and generates a full-length motion-graphics track — animated text, numbers, charts — timed to every line, which you drop straight onto your timeline in CapCut, Premiere or Final Cut.

Already have the .srt? You're one step from the visuals.

Upload subtitles, get a synced b-roll track back. Free credits to start, no card.

Start free

Frequently asked questions

How long does it take to transcribe a video by hand?
Roughly four times the runtime for a clean transcript — about 40 minutes of typing for a 10-minute video, more if there are multiple speakers or technical terms. Speech-to-text plus a cleanup pass usually cuts that to under 10 minutes.
What's the difference between a transcript and subtitles?
A transcript is the plain text of what was said. Subtitles are that text split into short lines, each stamped with a start and end time — usually stored as an .srt or .vtt file. Every subtitle file contains a transcript; a transcript on its own has no timings.
Can I get a transcript from a YouTube video?
Yes. On desktop, open the video, click the three dots below it and choose "Show transcript". YouTube's auto-captions are generated by speech recognition, so expect errors in names, jargon and numbers — read it through before using it for anything.
What file format should a transcript be in?
Plain text is fine if you only need the words. If you plan to do anything time-based with it — captions, clip selection, or generating visuals — export .srt instead, because it preserves the timing of every line.
How do I clean up an auto-generated transcript?
Fix proper nouns and numbers first, since those are where recognition fails most and where errors cost the most. Then add punctuation and paragraph breaks, and remove filler words only if the transcript is for reading rather than for captions.

Related: How to write a video script · How to add b-roll to videos automatically · What is b-roll? A-roll vs b-roll