YouTube's caption panel vs the transcript, same 60 seconds

Sample coming soon

Sample coming soon

The left column is the text YouTube's transcript panel shows for this video: the State Department uploaded its own caption file, so it has punctuation and the spokesperson's name. On a video whose channel uploaded no captions the panel shows auto captions: lowercase, no punctuation, no names. The right column is the same minute from the Transcription, with a speaker label and a timestamp on every paragraph.

What is an interview transcript

An interview transcript is the written record of what each person said, with a speaker label on every turn and a timestamp you can check against the video. A verbatim transcript keeps fillers, repeats and false starts.

What you walk away with

A transcript with a speaker label on every turn and a timestamp on every paragraph, 1:42 on screen. Click a timestamp and the video plays from that second. The .txt export writes 00:01:42 --> 00:01:47 above each paragraph, the .docx puts the speaker label in bold at the start of each paragraph, and a translation keeps both. Copy the line with its speaker and export timestamp, (Speaker, 00:01:42: "..."), and a reader can check it in the video.

A sample hearing with one quoted line, built from real outputs, comes with the sample.

Rename the speakers once

Rename Speaker 1 to the chair once, free, in the editor, and the name is applied everywhere: the transcript, the translation and every export. NVivo needs labels that stay the same through the whole transcript.

Words the engine was less sure about are flagged

The engine marks the words it was less sure about. Switch the Confidence view on to see them, click one to fix it, and check it against the video with the timestamp on its paragraph. Names and numbers are where you look first.

Edit, find and replace all

A name spelled three different ways gets fixed once with replace-all. Edits save to your account and undo is one keystroke away, so a slip is never final.

Ways to get an interview into text

Ways to get an interview into text
FeatureSpeaker labelsTimestampsVerbatimExportVideo length
Manual typingYes, by handYes, by handYes, if you type every umAny format you type in4 to 6 hours of typing per hour of audio
YouTube auto captionsNoPer caption line, copy and paste onlyNo punctuation, so no way to tellNone, copy and pasteOnly where YouTube made captions
TranscriptionYes, on every turnYes, on every paragraphOptional, per video.txt and .docx, with labels and timestampsUp to 12 hours

How it works

  1. 1

    Choose Transcription

    For a verbatim transcript, open the optional settings and switch "Clean up filler words" off. Leave it on for a cleaned transcript.

  2. 2

    Check the flagged words and rename the speakers

    Switch the Confidence view on, click each flagged word and fix it. Rename each speaker once, at no cost. Fix a misspelled name everywhere with replace-all.

  3. 3

    Listen before you quote

    Click the timestamp on a paragraph and the video plays from that second. Write the timestamp next to the quote in your notes, so a reader can check it.

  4. 4

    Export .txt or .docx

    Tick Speakers and Timestamps and the .txt carries a timestamp range on every paragraph. The .docx puts the speaker label in bold at the start of each paragraph.

Things worth knowing

Qualitative research: NVivo and ATLAS.ti import

Coding software wants one speaker label per paragraph and a label that never changes spelling. The .txt and .docx exports do that: tick Speakers and each paragraph starts with the label, tick Timestamps and each paragraph gets a time range above it. Rename the speakers in the editor before you export, so "Speaker 1" is "Chair" in every paragraph of every export. Import one transcript first and check that the software splits it the way you code, because every project has its own settings. Do that check once, then run the rest of the interviews the same way. For interviews that are not on YouTube, upload the audio or video at filetotext.ai and export the same .txt or .docx.

Foreign-language hearings

A hearing in Spanish or German comes back in Spanish or German. To read it in English, open the Translate tab and pick English. The translation keeps the speaker labels and the timestamps, so a line in the English text still points at the second in the video where it was said. One language per run; pick another language and the tab shows that one. Quote from the original when the exact wording matters, and use the translation to find the moment. Click the timestamp on the English line and the video plays the original from that second, so you can check both against the speaker's own words.

Overlapping speech and crosstalk

Two people talking at once is the hardest part of any hearing. The engine assigns each word to one speaker, so when two people overlap, the words land with the louder voice and the other speaker's words can be missing or attached to the wrong turn. The engine also marks the words it was less sure about, and crosstalk is where most of those flags sit. Switch the Confidence view on, click the timestamp on the paragraph of each flagged word, listen, and fix the word or move the turn to the right speaker. For a heated exchange, check that stretch by ear before you quote either side.

The problem in their words

the fact that the transcript isn't close to an exact match to what was said - and the timestamps are incorrect - means it's very hard to trust the output.
simonw on Hacker News
transcribing a 40-minute interview takes, oh, 1 1/2 hours
Jeff Pearlman, The Yang Slinger

What a typical 3-hour hearing costs

Free

10 min

Your first 10 minutes are free. They cover the start of a 3-hour hearing; the rest is locked until you unlock it.

Just this video

...

One payment for a 3-hour hearing, 180 minutes. No subscription.

Creator

... / month

600 minutes every month, the smallest plan that covers a 3-hour hearing.

The Deep Dive

...

600 minutes, one payment, credits never expire.

See all plans

FAQ

Questions.
Answered.

Can't find what you're looking for?
Contact us.

Only when you ask for it. By default the engine cleans the text: it drops ums, ahs and repeats and keeps the meaning. For a verbatim transcript, open the optional settings before you start and switch "Clean up filler words" off. Every filler, repeat and false start then stays in, which is what a research protocol or a disputed quote needs.
Yes, through the .txt and .docx exports. Tick Speakers and every paragraph starts with the speaker label; tick Timestamps and every paragraph carries a time range. Rename the speakers in the editor first so the labels are consistent. Import one transcript, check how your software splits it, then run the rest the same way.
Every paragraph has a timestamp, 1:42 on screen and 00:01:42 in the export. Click it, listen, and copy the wording from the transcript. APA puts the timestamp in the in-text citation, for example (Committee, 2026, 1:42). For a hearing, add the speaker before the quote, as in (Speaker, 00:01:42: "quote"), so a reader can check it in the video.
It works. The limit is 12 hours per video. Anything over 10 minutes is processed in parts and merged before you see it, so you get one transcript with one set of timestamps. A long video takes longer to come back, so start it and return later. Read the long video post for what changed and how to plan the minutes. Long videos, up to 12 hours
Not from the summary. The summary is a few paragraphs under one heading, in the video's language, with no timestamps. Ask AI answers questions with clickable timestamps, but it reads only the first 24,000 characters of the transcript, about 25 to 35 minutes of speech. For a 3-hour hearing, use Find in the editor and the timestamp on every paragraph instead.
The transcript step uses no language model. The engine is ElevenLabs Scribe v2, which writes what it hears and flags the words it was less sure about, so a wrong word shows up as a flag, not as a made-up sentence. The summary and Ask AI do use a language model. Quote from the transcript, and listen to the timestamp first.
Yes. The transcript comes back in the language spoken. Open the Translate tab, pick English, and the translation keeps every speaker label and timestamp, so each English line still points at its second in the video. Quote the original where the exact wording matters. The translate page shows the side-by-side view on a real sample. Translate a YouTube video

Updated September 2026

Ready when you are

Paste the interview link

Get the transcriptFirst 10 minutes free. No credit card required.