FAQ
Everything we get asked, in one place: what you get back, which videos work, languages and translation, editing and exports, plans, privacy, and how to reach the same engine from your own code or from an AI assistant. Nothing here is a claim we cannot show you in the product.
Yes. Paste the playlist or channel link instead of a single video and you get a list to pick from, up to 25 videos at a time. Each one runs as its own job, so a video that fails does not hold up the rest, and one page follows all of them. They spend the same minutes as anything else you transcribe.
Copy the link of the video, paste it in the box at the top of this page and start. You sign in first, then pick what you want: a transcript, a subtitle file, or captions burned into the video. When it is ready you get the text with punctuation, a label for each speaker and a timestamp on every paragraph.
Yes, and it is free. You sign in with Google, Apple or email, no card needed. The account is what keeps your videos, your corrections and your downloads together, so you can come back to a transcript later. Every account starts with 10 free minutes to try the quality on your own videos.
Your account starts with 10 free minutes, once, so you can try the product before paying anything. They work on the site and through the API. If your video is longer than the minutes you have left, we transcribe the first part and show it to you, and the rest stays locked until you unlock that video, buy a credit pack or take a plan.
Public and unlisted videos work. Private videos, members-only videos, age-restricted videos, live streams and videos that were removed or blocked in your country cannot be opened by us, and you get a message that says which of those it is. You paste one video link at a time.
Up to 12 hours. Anything over 10 minutes is split into parts, worked on in parallel and put back together, so you still get one transcript with the timings running through it. A long video uses more of your minutes, but it is not treated differently in any other way.
Reach us through the contact page. Tell us the link of the video and what went wrong, and we can look the job up. Billing questions, refunds, a transcript that came out badly and feature ideas all go through the same form, so nothing gets lost in a mailbox.
The full text of the video: punctuation, capital letters, a label for each speaker and a timestamp on every paragraph. You land in an editor where the transcript follows the video as it plays, so you can check anything by listening to it. From there you copy the text or download it in the format you need.
An SRT and a WebVTT of the same subtitles, cut into cues that follow the speech. Both come from one pass over the audio, so their timings match. Drop the SRT on a timeline in Premiere, DaVinci or CapCut, or upload it to your own video in YouTube Studio; use the WebVTT for a web player.
An MP4 with the captions rendered into the picture, ready to upload anywhere. The words are part of the video, so they show up on every platform whether or not the viewer switches captions on. You also get the subtitle file of the same job, in case your editor wants that instead.
Running text with punctuation and capital letters, split into paragraphs. A new paragraph starts whenever the speaker changes, and after at most 30 seconds of one person talking. Every paragraph carries a timestamp you can click to jump the video to that moment, and a label for the person speaking.
Yes. We do not read the captions on YouTube at all. We take the audio of the video and listen to it, so a video with no captions, with automatic captions full of mistakes, or with captions switched off by the creator all work the same way.
It separates the voices while it listens and gives each one a label, starting at Speaker 1. You can click any name and rename it to the real person, which changes it everywhere in the transcript at once. If a line was given to the wrong speaker you can move that line to the right one.
Yes. By default we clean up filler words and stutters, because that reads better. Turn the clean-up off before you start and you get everything exactly as it was said, which is what you want for a quote, an interview or anything you need on the record. You can see what would be removed before you decide.
A short summary of the video in paragraphs, in the language that is spoken in it. It covers the topics and the main points. It is written text, not a list of bullets, and it carries no timestamps. It is there to tell you what a video is about before you read the whole transcript.
Yes. Open Edit and use find and replace to jump to every mention of a word, with the number of hits shown as you type. Replace all fixes a name or a term through the whole text in one go. On a 12-hour video that is the difference between reading and scrubbing.
Yes. Anything over 10 minutes is cut into parts, transcribed in parallel and merged, and what you get back is one transcript with continuous timings. You never stitch anything together yourself, and the speaker labels carry across the whole video.
Those are the words the model was least sure about, usually names, brands and technical terms. Switch the confidence view on and they are marked so you can jump to them, listen to that moment and correct the ones that are wrong, instead of rereading the whole text hunting for mistakes.
That is what most people do with it. Read instead of watching, search for the one thing you need, and copy the text into your notes app: the paragraphs survive the paste. Ask for a summary first if you want to know what a long video covers before you read it.
Yes. Every paragraph has its own timestamp, so you can quote a line and say exactly when it was said. Keep the timestamps in the export and the times travel with the text; turn them off and you get clean prose for a document.
Yes. We start a new paragraph on every speaker change and after at most 30 seconds, which keeps long monologues readable. In the editor you can split a paragraph where a new thought starts, or join two that belong together, and the timestamps move with them.
A transcript is the text to read, search and copy. Subtitles are timed lines made to run under a video, so they are cut into short cues that appear and disappear with the speech. Pick Subtitles when you want an SRT or WebVTT for an editor or for YouTube, and Transcription when you want the text itself.
Yes. You get an MP4 with the captions rendered into the picture, which is what social platforms need when nobody turns the sound on. Choose 720p, 1080p or 4K. Burning in costs 3 times the video length at 720p, 4 times at 1080p and 5 times at 4K, because the video is encoded again.
One phrase or sentence per cue, at most 42 characters on a line and never more than two lines, each cue on screen between 1 and 7 seconds. That is what makes subtitles readable at a glance and it is why a transcript chopped into equal blocks never looks right under a video.
Yes. Every cue is editable in the browser: retype the text and the start and end times stay exactly where they were. Reset a cue to put the original text back. You can also download the file, edit it in another tool and upload it again to keep everything in one place.
Automatic captions come without punctuation or capital letters, carry no speaker names, and you cannot download them as a proper subtitle file from someone else's video. Ours are written from the audio with punctuation and cue timings, and you get the file itself, for any public video.
Yes, in two ways. Use the subtitle file if your editor takes one, or have the captions burned into the picture, which is what those platforms need when the sound stays off. Burned captions can be raised clear of the interface that sits at the bottom of a vertical video.
Yes, from the first word to the last. The cues run to the end of the video, and a long one is handled in parts and merged, so the numbering and the timings stay continuous. What you download is the complete file, not a sample of the opening minutes.
3 times the video length at 720p, 4 times at 1080p and 5 times at 4K. The multiplier is there because the whole video is encoded again at that quality, not because the captions cost more. A one-minute clip in 1080p uses 4 minutes.
Yes. Pick the font, the text colour, the outline and its thickness, an optional background box, the size as a share of the video height, and where the captions sit. You see the style over a frame of your own video before you apply it, so you are not guessing.
Not if you use the raised position. It anchors the captions above the bottom of the frame, clear of the buttons and the caption bar that TikTok, Reels and Shorts draw over a vertical video. Bottom, middle and top positions are there as well for other formats.
Fix the cue in the editor and burn it again. Restyling and re-burning cost nothing extra as long as the job is less than 7 days old, because we still hold the video that was downloaded. After that the source is gone and you would start a new job.
The video is encoded again, so pick the quality that matches your source: 720p for a phone clip, 1080p for most uploads, 4K when the original is 4K and you want to keep it. Burning a low-quality source at 4K makes the file bigger without making it look better.
Yes. Japanese, Korean, Chinese, Arabic, Hindi and other scripts are rendered in a font that covers them, whichever font you picked for Latin text, so nothing turns into empty boxes or question marks. Right-to-left text is laid out correctly in the picture, and the captions are sized the same way as they are for Latin text.
90+ languages, and you do not have to tell us which one it is: we work it out from the audio. You can also set the language yourself when a video mixes two of them or when the detection gets it wrong. Accents and dialects are handled by the same model.
Yes. Open a finished transcript, choose a language and you get the whole thing translated, with the speaker labels and the timestamps kept in place. You do one language at a time and can switch to another whenever you like. Translating a transcript you already have does not use extra minutes.
Yes. Pick the extra languages before you start and you get a subtitle file for each of them, with the cue timings matching the original so they stay in sync. You can have up to 10 languages including the one in the video. Each extra language costs one more time the video length.
Yes. Choose the extra languages before you start and you get a separate file per language, with the cue numbers and timings identical across all of them. Up to 10 languages including the one spoken in the video. Each extra language costs one more time the video length.
Yes. Pick the languages before you start and you get one MP4 per language, each with its own captions rendered in, plus the subtitle files. Up to 10 languages including the one spoken in the video. Every extra language adds one more time the video length.
Yes. Click Edit and you can change any word, rename a speaker, split or join paragraphs, and find and replace a word through the whole text. Changes save by themselves and you can undo them. Play the video while you read and the transcript follows along, so you can check a word by listening to it.
Copy it in one go, or download it as TXT, DOCX, SRT and WebVTT. Two checkboxes decide whether the speaker names and the timestamps come along, so you can take clean prose into a document or keep the timings for reference. Subtitle files keep their cue times whatever you choose.
Yes. Create a share link and anyone with that link can read the transcript on a page of its own, without an account and without seeing anything else from yours. They can read and copy the text; they cannot change it. Delete the transcript and the link stops working.
TXT for plain text you paste anywhere, DOCX for a Word document with the speakers in bold. The SRT and WebVTT downloads here carry the transcript in a subtitle layout, which is fine for reference. For real subtitle cues that an editor expects, start a Subtitles job instead.
Yes, but start a Transcription job for it: that gives you the running text with speaker labels, timestamps and the option of a summary. A subtitle file is cut for reading under a video, so it makes for choppy prose if you paste it into a document.
Not here: this tool starts from a YouTube link. If the video lives on your computer, our sister tool FileToText takes the upload instead. For anything that is already on YouTube, public or unlisted, paste the link and pick the burned option.
Our benchmark measures 98.5% on clean read speech, 97.1% on harder audio, 94.3% on accented speech and 92.7% on talks: 95.7% across the set. Audio quality, accents and background noise move the number for any one video. We measure this every month on public speech collections that come with a human transcript to compare against, and we publish what came out, including the conditions where it scores worst.
Fix them in the editor: the words we were least sure about are marked, so you can jump straight to them, listen and correct. Music, crosstalk and heavy background noise are where any tool struggles. If the result is not worth what you paid, ask us for a refund and you get it.
Creator is $10 a month for 600 minutes, Creator+ is $20 a month for 1,800 minutes and Pro is $50 a month for 6,000 minutes every month. Pay for a year and you pay for 10 months and get 12 months of minutes. Plan minutes reset at the start of each month. If you need more in a busy month you can add credits, which sit on top of your plan.
Yes, in two ways. Credit packs are a one-off payment for a pot of minutes that never expires: The Deep Dive is $25 for 600 minutes, and there are bigger ones. Or pay for a single video: $3 for a video up to an hour, rising to $12 for the longest ones. Both work without a monthly plan.
Minutes that come with a monthly plan reset when the new month starts, so they do not pile up. Minutes you bought as credits are yours until you use them: they never expire. Unlocking one video separately does not touch either pot.
Open the billing page in your account and cancel there, or ask us through the contact page and we do it for you. There is no notice period and nothing to negotiate. Your transcripts stay in your account after you cancel, so you can still read and download what you already made.
Transcribing stops until your minutes reset, and nothing is charged. You can wait, buy a credit pack, or move up a plan. There is also extra usage, which is off until you switch it on yourself in Billing: with it on, anything past your included minutes costs $1 an hour on Creator, billed in whole hours. Each plan has its own rate and it lands on your next invoice.
Yes, no questions asked. If the result was not what you expected or you bought the wrong thing, ask through the contact page and we refund the purchase. We would rather give the money back than have you fight with a tool that did not do the job.
A transcript costs the length of the video: a 40-minute podcast uses 40 minutes. Asking for a summary or a translation of a transcript you already have costs nothing extra. Your first 10 minutes are free, so you can try a video before deciding to pay for the rest.
One time the length of the video, the same as a transcript, plus one more time the length for every extra language you pick. Burning the captions into the picture costs more, because the video is encoded again. There is no separate charge for the editor or the downloads.
You, through your account. Nobody else can open them, and they are not listed or indexed anywhere. The one exception is a share link: if you create one and pass it on, whoever holds it can read that transcript. Delete the transcript and the link dies with it.
The audio and video we downloaded are deleted after 7 days. The transcript itself stays in your account until you delete it, so it is there when you come back. Delete a transcript and it goes for good, together with its subtitle files and any burned-in video.
No. Your audio goes to the speech-to-text engine we use to turn it into words, and the text goes to the model that writes summaries and translations when you ask for one. It is processed to give you the result and for nothing else. Payments run through Stripe, which never sees your video.
Yes. Send a video link, poll the job, read back the transcript, subtitle files or a captioned video. API and MCP are included on every plan. Your 10 free minutes work through the API too. You generate one token in your account and send it as a bearer token. The reference lists every endpoint with the request and the response you should expect.
Yes, through our MCP server. Add it once as a connector in Claude, ChatGPT, Cursor or Cline, sign in with your account, and then ask your assistant for the transcript or the subtitles of a video in plain language. It uses the same minutes as the site.
Sign in, open the API page in your account and generate one. The token is shown once, so copy it there and then: we keep only a hash of it. One token is active per account and you can revoke it whenever you want, which immediately stops anything using it. Tokens do not expire on their own.
Put the token in an Authorization header as a bearer token, then post a YouTube link to the transcribe endpoint. You get back a job id. Poll the transcription endpoint with that id until the state says done, and read the text off the response. The reference shows the exact request and response for each step.
Three things, one endpoint each: a timestamped transcript with speaker tags, subtitles as SRT and WebVTT, or a video with the captions burned into the picture. Each takes a YouTube link and the options you would pick on the site, such as the languages you want or the quality to burn at.
No. Creating a job answers straight away with an id, because transcribing takes as long as the video needs. You poll for the result. A job moves through states you can show your own users, from waiting and downloading to processing and done, and a failed job carries the reason.
API and MCP are included on every plan. Your 10 free minutes work through the API too. A job through the API uses minutes exactly as the same job on the site does, out of the same balance, so there is no separate developer plan or per-request fee to think about. The 10 free minutes every new account starts with are enough to build against before you pay anything.
It is the same product, reachable from an AI assistant instead of from your code. You add it once as a connector, and after that you can ask for the transcript or the subtitles of a video in plain language and the assistant fetches them. Nothing is copied and pasted between tabs.
Claude, ChatGPT, Cursor and Cline, and anything else that speaks MCP over a remote connection with OAuth. Add the server URL from the MCP page as a custom connector. If your client asks for a client id or a client secret, leave both blank: the server hands them out itself.
Your assistant opens a browser tab the first time you use it and you sign in with the account you already have. There is no token to paste and no key to store in a config. You stay connected for a year of use, and reconnecting is the same two clicks.
Paste a link and ask for a plain transcript, a verbatim one that keeps every filler word, a subtitle file in the language you name, or a video with the captions burned in at a quality you pick. It can also look up a job you started earlier and read the result back to you.
Use the API when a program of yours needs transcripts: a script, a backend, an automation. Use the MCP server when you are the one asking, inside an assistant you already work in. They run on the same account, the same balance and the same engine, so you can start with one and add the other.
Tell us the link of the video and what went wrong and we can look the job up. Building something instead? The API reference lists every endpoint with its request and response, and the MCP page walks through connecting an assistant.