Audio to Text · Recordings and video
Upload audio or video and the language is detected for you. When it finishes, copy the text, download a TXT, or export SRT subtitles. One file up to 2 hours or 50 MB.
No transcripts yet
Your recent audio-to-text projects will appear here.
Drag in an audio or video file — AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV are accepted, up to 2 hours or 50 MB. Format, size, and length are checked before the upload starts.
Optional, and worth the few seconds for anything a general model has never heard: personal names, brands, acronyms, in-house jargon. Up to 50 terms, remembered on this device for your next upload.
Transcription runs on our servers, so you can close the tab and come back to it. The finished page shows duration and word count above the full text, ready to copy, download as UTF-8 TXT, or export as SRT subtitles.
Get a quotable record of a recorded interview, then search it for the line you half remember instead of scrubbing back through the audio.
Turn a call recording into one continuous transcript you can paste into project notes and pass to whoever missed the meeting.
Read a class or conference talk back at your own pace, and keep the terminology-heavy stretches as text you can annotate.
Export SRT built from the sentence timings and load it straight into your editor instead of typing captions by hand.
Speech recognition listens through a recording and writes out what was said as text. It replaces typing the whole thing up by hand, so an hour of audio becomes a document you can read and search in a fraction of the time.
Seventeen audio and video formats: AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV. Video files go in as they are — there is no need to extract the audio track first.
One file at a time, up to 2 hours long and 50 MB. Both limits are checked in the browser first, so an oversized file is caught immediately instead of after a long upload.
No. The spoken language is detected from the recording itself, which also means the file does not have to match the language of the interface.
Yes. Each sentence is stored with its start and end time, so the same result exports as an SRT subtitle file alongside the UTF-8 TXT transcript. If a recording yields no usable sentence timings, the SRT option is disabled rather than guessed at.
One credit for every started minute of the original recording, so a 90-second file costs 2 credits. The estimate appears on the submit button before you start, and nothing is charged if the transcription fails.
That depends on the recording. Clear speech, one person talking at a time, and low background noise all transcribe better. Names, brands, and specialist terms are the usual weak spot, which is exactly what the reference vocabulary field is there to fix.