Home
Voice Library
Audio
AI VoiceNew
Text to Speech
AI Podcast
Voice Cloning
Visual
AI Image
AI Video
Explainer Video
Slides
Blog
English
Sign In

ListenHub · Voice Library · 12 languages

AI Voice Library

Hear the voice before you write the script

ListenHub speaks with 909 official voices across 12 languages, and every one of them plays right here. Find the one that fits, then take it straight into Text to Speech.

  • 909 official voices
  • 12 languages · English 142 · Chinese 491 · Japanese 35
  • Free to play, no account
  • Filter by scene and style
  • One click into the editor
English142Chinese491Japanese35Spanish65Portuguese86French10German5Turkish2Korean53Italian5Thai7Vietnamese8
Any voiceFemaleMale
Scene

1–7 of 7 voices

Valeria

Female · Thai

"Valeria" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Zara

Female · Thai

"Zara" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Valentina

Female · Thai

"Valentina" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Mildred

Female · Thai

"Mildred" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Lydia

Female · Thai

"Lydia" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Iris

Female · Thai

"Iris" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

Phoebe

Female · Thai

"Phoebe" is an official ListenHub Thai female voice, suited to general audio creation, with a natural delivery.

  • General
Read text aloudSkills docs

How to find a voice here

  1. 1

    Start from the language

    Each of the 12 languages has its own voices: the Chinese set is not the English one translated, it is 491 different voices. Pick the language you will publish in first.

  2. 2

    Narrow by what you are making

    Scene tags say what a voice is for: podcast, audiobook, documentary, kids. Style tags say how it delivers. Stack two of them and a few hundred voices become a handful.

  3. 3

    Play a few back to back

    Every card plays the same demo line, so what you are comparing is the voice rather than the script it happens to be reading.

  4. 4

    Take it into a tool

    Read text aloud opens Text to Speech with that voice already selected. The existing three-language voices also keep their Podcast and AI Voice shortcuts.

Use these voices from your agent

Every voice here has a speaker ID, and that is what ListenHub Skills, the MCP server, the CLI and the API all use to pick a voice. Copy one off a card, paste it to your assistant, and the audio it makes comes back in that voice.

  1. 1

    Install the skills

    One command inside your project. Claude Code, Cursor, Windsurf and OpenCode all read Agent Skills.

    npx skills add marswaveai/skills
  2. 2

    Set your API key once

    Create a key starting with lh_sk_. Your assistant asks for it the first time it needs one and keeps it after that.

    Create an API key
  3. 3

    Copy a voice, paste the prompt

    Press Copy agent prompt on any card above. The sentence arrives with the voice name and its speaker ID already in it — add whatever you want read, and send.

What lands on your clipboard

Use the ListenHub voice "Valeria" (speakerId: doubao-official-th_female_bv568_angry_uranus_bigtts) to turn the following into audio:

Any card gives you the same sentence with its own voice in it.

Other ways in

What people come here to cast

A two-host show

Duo episodes need two voices that sound like different people, not the same person twice. Play them side by side before you commit to the pair.

Something long to listen to

An audiobook or a long read lives or dies on whether the voice is still pleasant forty minutes in. Steady beats striking here.

Explainers and course narration

Teaching material wants clear over characterful — the documentary and education tags are where those voices sit.

Dubbing and characters

Video dubbing and character work need range rather than neutrality. Filter by the character and multi-emotion styles to find voices that act.

Voice library FAQ

Do I need an account to listen?

No. Every demo on this page plays without signing in. An account only comes into it when you want to generate audio of your own.

How many voices are there?

909 today across 12 languages, including 142 English, 491 Chinese and 35 Japanese voices. The page counts what is actually published rather than quoting a stale hardcode.

Are the voices free to use?

Listening is free. Generating audio spends credits from your plan, and what it costs depends on the format and length. A few voices are marked Subscription: you can audition them here, but generating with them needs a plan.

What do the scene and style tags mean?

Scene is what a voice is for — podcast, audiobook, news, meditation. Style is how it speaks — neutral, conversational, character, multi-emotion. They come from the same vocabulary the voice picker uses inside the app, so a voice you find here behaves the way you expect once you are in the editor.

Can I use two different voices in one episode?

Yes. A duo podcast is written for two hosts and you choose both of them, which is exactly why it helps to audition them next to each other first.

Can I use my own voice instead?

Yes, through voice cloning — 25 to 35 seconds of clean audio and your voice joins your own picker. Cloning supports the same 12 languages as Text to Speech. Cloned voices stay private to your account and never appear in this public library.

What do the tags and descriptions tell me?

Every published profile carries scene or style tags and a written description. Use the tags to narrow the list, read the description for intent, then press play to judge the voice itself.

Where to go instead

  • AI Podcast Generator

    When you want two hosts talking it through rather than one voice reading it out.

  • Text to Speech Generator

    When the script is already written and you just need it spoken.

  • AI Voice Generator

    When the delivery matters — several voices in one take, or a line that has to land with real emotion.

  • AI Voice Cloning

    When none of the library voices should be the one speaking — record 30 seconds and use your own.

Command line→

listenhub tts create --speaker-id <id>

Podcasts, explainer videos and slides take the same flag.

  • MCP server→

    Voices and generation as tool calls inside Claude Desktop, Cursor and Zed.

  • OpenAPI→

    Calling it yourself: the id goes in scripts[].speakerId, or in voice.

  • The whole catalogue as text→

    Every voice, id and tag in one plain-text fetch, with no pages to walk.

    https://listenhub.ai/voices.txt