Paste the link of an interview, hearing or press briefing and get every turn with a speaker label and a timestamp on every paragraph.
Trusted by 42,800+ creators
Sample coming soon
Sample coming soon
The left column is the text YouTube's transcript panel shows for this video: the State Department uploaded its own caption file, so it has punctuation and the spokesperson's name. On a video whose channel uploaded no captions the panel shows auto captions: lowercase, no punctuation, no names. The right column is the same minute from the Transcription, with a speaker label and a timestamp on every paragraph.
An interview transcript is the written record of what each person said, with a speaker label on every turn and a timestamp you can check against the video. A verbatim transcript keeps fillers, repeats and false starts.
A sample hearing with one quoted line, built from real outputs, comes with the sample.
Rename Speaker 1 to the chair once, free, in the editor, and the name is applied everywhere: the transcript, the translation and every export. NVivo needs labels that stay the same through the whole transcript.
The engine marks the words it was less sure about. Switch the Confidence view on to see them, click one to fix it, and check it against the video with the timestamp on its paragraph. Names and numbers are where you look first.
A name spelled three different ways gets fixed once with replace-all. Edits save to your account and undo is one keystroke away, so a slip is never final.
| Feature | Speaker labels | Timestamps | Verbatim | Export | Video length |
|---|---|---|---|---|---|
| Manual typing | Yes, by hand | Yes, by hand | Yes, if you type every um | Any format you type in | 4 to 6 hours of typing per hour of audio |
| YouTube auto captions | No | Per caption line, copy and paste only | No punctuation, so no way to tell | None, copy and paste | Only where YouTube made captions |
| Transcription | Yes, on every turn | Yes, on every paragraph | Optional, per video | .txt and .docx, with labels and timestamps | Up to 12 hours |
For a verbatim transcript, open the optional settings and switch "Clean up filler words" off. Leave it on for a cleaned transcript.
Switch the Confidence view on, click each flagged word and fix it. Rename each speaker once, at no cost. Fix a misspelled name everywhere with replace-all.
Click the timestamp on a paragraph and the video plays from that second. Write the timestamp next to the quote in your notes, so a reader can check it.
Tick Speakers and Timestamps and the .txt carries a timestamp range on every paragraph. The .docx puts the speaker label in bold at the start of each paragraph.
Coding software wants one speaker label per paragraph and a label that never changes spelling. The .txt and .docx exports do that: tick Speakers and each paragraph starts with the label, tick Timestamps and each paragraph gets a time range above it. Rename the speakers in the editor before you export, so "Speaker 1" is "Chair" in every paragraph of every export. Import one transcript first and check that the software splits it the way you code, because every project has its own settings. Do that check once, then run the rest of the interviews the same way. For interviews that are not on YouTube, upload the audio or video at filetotext.ai and export the same .txt or .docx.
A hearing in Spanish or German comes back in Spanish or German. To read it in English, open the Translate tab and pick English. The translation keeps the speaker labels and the timestamps, so a line in the English text still points at the second in the video where it was said. One language per run; pick another language and the tab shows that one. Quote from the original when the exact wording matters, and use the translation to find the moment. Click the timestamp on the English line and the video plays the original from that second, so you can check both against the speaker's own words.
Two people talking at once is the hardest part of any hearing. The engine assigns each word to one speaker, so when two people overlap, the words land with the louder voice and the other speaker's words can be missing or attached to the wrong turn. The engine also marks the words it was less sure about, and crosstalk is where most of those flags sit. Switch the Confidence view on, click the timestamp on the paragraph of each flagged word, listen, and fix the word or move the turn to the right speaker. For a heated exchange, check that stretch by ear before you quote either side.
the fact that the transcript isn't close to an exact match to what was said - and the timestamps are incorrect - means it's very hard to trust the output.
transcribing a 40-minute interview takes, oh, 1 1/2 hours
10 min
Your first 10 minutes are free. They cover the start of a 3-hour hearing; the rest is locked until you unlock it.
...
One payment for a 3-hour hearing, 180 minutes. No subscription.
... / month
600 minutes every month, the smallest plan that covers a 3-hour hearing.
...
600 minutes, one payment, credits never expire.
Updated September 2026
Ready when you are