Home
Audio6
Video9
Image2
All Tools
Assets
Voice Library
Developer Docs
Blog
English
Pricing
Sign In
Download App
AudioAudio to Text
AI PodcastText to SpeechAudio to TextVoice CloningAI VoiceMulti-Speaker Voiceover

Audio to Text · Recordings and video

Audio to Text Converter

Turn a recording into a written transcript

Upload audio or video and the language is detected for you. When it finishes, copy the text, download a TXT, or export SRT subtitles. One file up to 2 hours or 50 MB.

  • Automatic language detection
  • 17 audio and video formats
  • Reference vocabulary for names and jargon
  • Export TXT or SRT subtitles

No transcripts yet

Your recent audio-to-text projects will appear here.

How to convert audio to text

  1. 1

    Upload one recording

    Drag in an audio or video file — AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV are accepted, up to 2 hours or 50 MB. Format, size, and length are checked before the upload starts.

  2. 2

    Add reference vocabulary

    Optional, and worth the few seconds for anything a general model has never heard: personal names, brands, acronyms, in-house jargon. Up to 50 terms, remembered on this device for your next upload.

  3. 3

    Copy, download, or caption

    Transcription runs on our servers, so you can close the tab and come back to it. The finished page shows duration and word count above the full text, ready to copy, download as UTF-8 TXT, or export as SRT subtitles.

Ways to use audio transcription

Interview transcripts

Get a quotable record of a recorded interview, then search it for the line you half remember instead of scrubbing back through the audio.

Meeting minutes

Turn a call recording into one continuous transcript you can paste into project notes and pass to whoever missed the meeting.

Lectures and study notes

Read a class or conference talk back at your own pace, and keep the terminology-heavy stretches as text you can annotate.

Subtitles for video

Export SRT built from the sentence timings and load it straight into your editor instead of typing captions by hand.

Audio to Text FAQ

What is audio-to-text transcription?

Speech recognition listens through a recording and writes out what was said as text. It replaces typing the whole thing up by hand, so an hour of audio becomes a document you can read and search in a fraction of the time.

Which file formats are supported?

Seventeen audio and video formats: AAC, AMR, AVI, FLAC, FLV, M4A, MKV, MOV, MP3, MP4, MPEG, OGG, OPUS, WAV, WEBM, WMA, and WMV. Video files go in as they are — there is no need to extract the audio track first.

How long or how large can a recording be?

One file at a time, up to 2 hours long and 50 MB. Both limits are checked in the browser first, so an oversized file is caught immediately instead of after a long upload.

Do I have to choose the language?

No. The spoken language is detected from the recording itself, which also means the file does not have to match the language of the interface.

Can I get subtitles as well as plain text?

Yes. Each sentence is stored with its start and end time, so the same result exports as an SRT subtitle file alongside the UTF-8 TXT transcript. If a recording yields no usable sentence timings, the SRT option is disabled rather than guessed at.

How are credits calculated?

One credit for every started minute of the original recording, so a 90-second file costs 2 credits. The estimate appears on the submit button before you start, and nothing is charged if the transcription fails.

How accurate is the transcript?

That depends on the recording. Clear speech, one person talking at a time, and low background noise all transcribe better. Names, brands, and specialist terms are the usual weak spot, which is exactly what the reference vocabulary field is there to fix.

Where to go instead

  • AI Podcast Generator

    When you want two hosts talking it through rather than one voice reading it out.

  • Text to Speech Generator

    When the script is already written and you just need it spoken.