Comparisons

The best text-to-speech tools in 2026, compared

Eleven tools compared on price, voice quality and how well an AI agent can drive them, plus a full account of what ListenHub does — API, MCP and Agent Skills included — and where we are the wrong choice.

ListenHub TeamPublished
A dark purple and magenta gradient title card labeled COMPARISON, headlined "The best text-to-speech tools, compared", with the ListenHub wordmark above listenhub.ai

ListenHub publishes this post. We make text-to-speech, AI podcasts and voice cloning, and we sell subscriptions, so we have an obvious stake in what you conclude. Ignoring that would be the fastest way to lose you.

This article does two jobs, and it is worth knowing which is which. The first is a straight comparison of eleven text-to-speech products, grouped by the job you are hiring them for, with every price read off the vendor's own pricing page rather than a listicle. The second is a detailed account of what ListenHub does — including the developer and agent surface, which is the part of our product most people never see — and an equally detailed account of where we are the wrong choice. We award nobody a score out of ten, least of all ourselves, and one section near the bottom exists only to send you to a competitor.

The short answer

If you only read one section, read this one.

  • One voice reading a long script beautifullyElevenLabs, $6/month.
  • An API with volume and a budgetAmazon Polly or Google Cloud, $4 per million characters.
  • Real-time voice agentsDeepgram or Cartesia.
  • Compliance-reviewed corporate e-learningWellSaid Labs.
  • Narration inside a video timelineSpeechify Studio.
  • A document turned into a private listen, for freeNotebookLM.
  • An agent or codebase that produces finished audio and video on its ownListenHub. One key covers podcasts, speech, cloning, music, images, slides and video, reachable over REST, an MCP server, Agent Skills, an SDK and a CLI.
  • A finished, publishable multi-voice program from a documentListenHub. These last two are the jobs where we think we are the right answer; the sections below explain why, and where we still fall short.

At a glance

Prices checked July 28, 2026. Units differ by vendor and are stated per row rather than converted into a fake common currency.

ToolBest forFree tierPaid entryVoice cloningAPI / agent access
ElevenLabsLong-form narration quality10k credits/mo, no commercial use$6/moYes, from $6REST, SDKs, official MCP server
Amazon PollyCheap API at volume5M chars/mo, ongoingPay as you goNoAWS SDK, IAM
Google Cloud TTSWidest voice range via API4M chars/mo, ongoingPay as you goNoGCP SDK
OpenAIOne vendor, one keyNonePay as you goEnterprise onlyREST, SDKs
DeepgramReal-time agents$200 credit, oncePay as you goNoREST, streaming SDKs
CartesiaCheap low-latency streaming20k credits/mo, no commercial use$5/moYes, from $5REST, streaming SDKs
Azure AI SpeechEnterprise Microsoft stacks500k neural chars/moPay as you goYes, gatedAzure SDK
WellSaid LabsCorporate e-learning3 download min/mo, no commercial use$19/moNoREST
Speechify StudioVideo-timeline voiceover600 credits, no commercial use$19/moYes, from $19REST
NotebookLMFree personal audio overviews3 generations/dayGoogle AI subscriptionNoNone
ListenHubFinished programs, and agents that build themMonthly plus daily allowance$12/mo, or $9/mo annuallyYes, all paid plansREST, MCP server, Agent Skills, SDK, CLI — any account

Billing units, since nobody prices this the same way: Google, AWS and OpenAI's legacy models bill per character, ElevenLabs, Cartesia, Speechify and ListenHub in credits, WellSaid in downloaded minutes, OpenAI's newer models in tokens. Those units are not interchangeable, and any comparison that adds them into one league table is making it up.

The last column is the one most comparisons omit, and in 2026 it decides more purchases than voice quality does. Everything here has a REST API. What differs is whether an AI agent can drive it without you writing an integration first, and what one call actually gives you back — audio bytes, or a finished piece of content.

How we compared, and what we could not test

What we compared, and when. Published pricing, free-tier limits and commercial-use terms, read off each vendor's own pricing page on July 28, 2026 and cross-checked against their docs — plus the capability questions that decide real purchases: does your tier allow you to publish, does it include voice cloning, is there a pronunciation lexicon, what unit are you billed in. TTS pricing moves constantly, so treat every number here as a starting point and re-check before you put a card in.

Why there are no scores. A rating out of ten implies a scored panel we did not run. Where we say a tool sounds better, that is our opinion from using these products and from building our own. Treat those lines as a shortlist to test yourself, and use the failure modes in the next section as your scoring sheet. Twenty minutes with your own worst paragraph beats anyone's leaderboard, including ours.

What we could not verify. Some vendors render prices in JavaScript, leaving no stable published figure to quote. Murf is described below without numbers for that reason, and Microsoft publishes its free tier as text but loads paid rates client-side, so we quote only the free allowance. PlayHT, LOVO and MiniMax were dropped entirely: no first-party price would load, and the third-party figures contradict each other and are visibly recycled from earlier years. We would rather cover eleven tools we can source than fifteen we cannot.

What we could not test. Enterprise-gated features — OpenAI's custom voices require a sales conversation, so we cannot tell you how they sound. We also did not benchmark latency under production load, because a single-request timing tells you nothing about a thousand concurrent streams.

What actually makes text-to-speech sound wrong

Nearly every tool here clears the old bar — the voices are no longer robotic. The failures that remain are subtler, and they are the ones your audience notices without being able to name. This is the part worth training your ear on before you spend money.

Flat prosody. English carries meaning in stress. "I didn't say she took the money" has about seven meanings depending on which word you lean on. A model that picks the statistically average stress pattern will read your sentence correctly and mean the wrong thing. Listen for contrast: does the voice emphasize the word that is actually new information?

Text written for the eye, read literally. This is the big one, and it is not a voice problem at all. Your document says "e.g.", "Fig. 3", "$1.2M", "2026-07-28", "(see below)". A human reading aloud silently rewrites all of that — "for example", "figure three", "one point two million dollars". Most engines just voice the glyphs. Bullet fragments are worse: written lists have no verbs, so they come out as a stack of disconnected noun phrases. We built a written-to-spoken rewrite step specifically because this failure survives every improvement in voice quality.

Names and jargon, wrong every single time. Product names, non-English surnames, acronyms that are sometimes spelled out and sometimes said as a word — "SQL", "Nginx", "Xiaomi", "José". The consistency is what makes it lethal: it will mispronounce your company name identically in all forty episodes. Check whether a vendor has a pronunciation lexicon you can edit, and whether it applies workspace-wide or has to be re-entered per project.

No turn-taking. Two voices rendered separately and stitched end to end is not a conversation. Real dialogue overlaps slightly, carries backchannels, and — most tellingly — the second speaker's opening pitch responds to where the first speaker's pitch landed. Splice two independent renders together and you get two monologues in a trench coat.

Uniform pacing. Human readers accelerate through a list and slow into a conclusion, and they breathe at paragraph boundaries rather than at commas. Engines that apply one fixed gap between sentences produce audio that is exhausting past about two minutes — and that is very hard to hear in a fifteen-second demo.

Which raises the demo trap. Vendor sample sentences are chosen because they work. Judge nothing on them.

Natural narration for video and courses

The largest group of buyers: you have a script, you need a voice reading it, and you need the rights to publish the result.

ElevenLabs — best for long-form narration

Free $0/mo, 10,000 credits, no commercial license, no cloning · Paid entry $6/mo Starter · Unit credits

Starter at $6/month is the real entry point, not the free tier: 30,000 credits, a commercial license, and instant voice cloning. Creator is $22/month for 121,000 credits and professional voice cloning; Pro is $99/month for 600,000. On long-form narration it holds prosody better than anything else we have used — it is the tool least likely to lose the thread of a sentence three paragraphs into a chapter.

Good: the strongest long-form prosody on this list; instant cloning at the cheapest paid tier; strong multilingual output from a single cloned voice; a genuine ecosystem of docs and integrations.

Not so good: the free tier grants no commercial license and no cloning, which trips up people who evaluate on free and then publish; credits are hard to convert to minutes in your head; costs climb quickly past the Creator tier.

Bottom line: if the job is one voice reading a long script and it has to sound right, buy Starter and stop shopping.

WellSaid Labs — best for corporate e-learning

Free 3 download minutes/mo, no commercial rights · Paid entry $19/mo, or $120/yr · Unit downloaded minutes

WellSaid sells to corporate and e-learning teams and prices in downloaded minutes, which is refreshingly easy to budget against. Starter is $19/month, or $120/year annually, for 20 download minutes a month; Pro is $49/month, or $396/year, for 180. The voices are fewer and more conservative than ElevenLabs — exactly what a compliance-reviewed training module wants.

Good: minutes are a unit a finance team understands; conservative, consistent voices that survive legal review; annual pricing that is a real discount rather than a rounding trick.

Not so good: a small catalogue by 2026 standards; no voice cloning; 20 minutes a month is thin if you produce weekly.

Bottom line: the safe procurement answer for training content, and the easiest bill on this list to forecast.

Speechify Studio — best for voiceover inside a video edit

Free 600 credits, no commercial rights · Paid entry $19/mo Starter · Unit credits, 1 per second of voiceover

Speechify Studio is built around a video editing timeline rather than a text box. Starter is $19/month for 7,200 credits with voice cloning and commercial rights; Creator is $49/month for 28,800. Voiceover costs 1 credit per second, so Starter is two hours of finished narration — a rare credit system you can convert in your head without a spreadsheet.

Good: a credit-per-second rate you can actually reason about; voiceover and video editing in one place; cloning and commercial rights at the entry tier.

Not so good: you are paying for an editor you may not want; narration quality is a step below ElevenLabs on long scripts; the free tier is a demo.

Bottom line: the right pick when the audio is one track in a video project, not the deliverable.

Murf — best for collaborative studio work

Free not quoted · Paid entry not quoted · Unit not quoted

Murf belongs in this group, particularly for teams who want a collaborative studio with shared projects. We are not quoting prices because its pricing page renders client-side and we could not read a figure we would stand behind. Everything else here has a number attached; treat the absence as a gap in our sourcing rather than a verdict on the product.

Bottom line: worth a look for team workflows, but get the current pricing from Murf directly.

An API you can scale

Here the question is not which voice is prettiest. Four things decide it instead.

Unit economics. What you pay per character, per second or per job, and whether that curve stays sane at your volume. The clouds win outright on raw synthesis: $4 per million characters is a floor no subscription can go under.

Latency and shape. Whether you need a streaming socket for a conversational turn or a fire-and-forget job for a ten-minute render. These are different products, and a vendor optimised for one is usually mediocre at the other.

What one call returns. This is the axis people skip. Most of this section returns audio bytes for text you supply — you own the script, the segmentation, the multi-voice assembly, the subtitle timing and the muxing. A smaller group returns a finished artefact. That difference is worth more engineering time than any price gap on this page.

Whether an agent can drive it unaided. In 2026 a lot of these calls are made by a coding agent rather than a developer at a keyboard. What matters then is whether the vendor ships a tool surface an agent can discover and call — an MCP server, an Agent Skill, a typed SDK — or whether someone has to write and maintain the integration first.

Also worth knowing before you commit: vendor durability. Google, AWS, Microsoft and OpenAI are not going anywhere. The rest of this list is younger, and pricing on newer entrants has moved more than once.

Google Cloud Text-to-Speech — best for range

Free 4M chars/mo Standard and WaveNet, 1M Chirp 3 HD, recurring · Paid $4–$160 per 1M chars · Unit characters

The widest useful range on this list. Standard and WaveNet are $4 per million characters after the recurring free allowance; Neural2 is $16 after a free 1 million; Chirp 3 HD, the current flagship, is $30 after a free 1 million; Studio is $160. The newer Gemini-TTS models bill in tokens and carry no free allowance. One gotcha: you must enable billing before you can touch even the free characters.

Good: recurring free tier large enough to run a real prototype at zero cost; a voice class for every budget; dozens of languages with locale variants.

Not so good: the price ladder from $4 to $160 punishes anyone who picks a voice class without reading; you write code, there is no studio; the token-billed Gemini models break comparison with everything else.

Bottom line: the default API choice when you want one vendor to cover both cheap bulk and a flagship voice.

Amazon Polly — best for the cheapest permanent free tier

Free 5M chars/mo Standard, ongoing with no 12-month cliff · Paid $4–$100 per 1M chars · Unit characters

The price leader at the low end and the best free tier in the business. Standard voices are $4 per million characters, Neural $16, Generative $30, Long-Form $100. The 5 million Standard characters a month do not expire after a year, which is rare enough to be the headline.

Good: the only genuinely permanent free allowance at scale; boring in the way infrastructure should be; predictable per-character billing.

Not so good: the better voice classes are time-limited on free; no cloning; Standard voices sound dated next to the 2026 flagships.

Bottom line: if free-forever matters and you can write code, start here.

OpenAI — best for teams already on OpenAI

Free none · Paid tts-1 $15, tts-1-hd $30 per 1M chars · Unit characters, or tokens on newer models

The obvious choice if your stack is already there and you want one vendor and one key. The legacy models are easy to reason about: tts-1 at $15 per million characters, tts-1-hd at $30. The newer gpt-4o-mini-tts bills in tokens rather than characters, so you cannot compare it to a per-character vendor until you have measured your own audio token output.

Good: one key, one bill, one SDK if you are already there; strong instruction-following on tone; good docs.

Not so good: no free tier at all — you pay from the first request; token billing on the newer models is genuinely hard to forecast; custom voices are gated behind a sales conversation. OpenAI's usage policies also require you to disclose that the voice is AI-generated.

Bottom line: convenience is the feature. Nobody switches to OpenAI for the voices alone.

Deepgram — best for real-time agents

Free $200 credit, one time, no card · Paid Aura-2 $0.030 per 1k chars, Aura-1 $0.0150 · Unit characters

Built for real-time and agent workloads, where latency matters more than beauty. Signup includes $200 in credit with no card — the most generous no-commitment trial here, though one-time rather than recurring.

Good: the trial is large enough to ship a prototype on; latency built for conversational turns; speech-to-text from the same vendor.

Not so good: the $200 does not come back; a narrow voice catalogue; not aimed at long-form narration.

Bottom line: the right answer for a voice agent, the wrong one for an audiobook.

Cartesia — best for cheap streaming

Free 20k credits/mo, no commercial use · Paid entry $5/mo Pro · Unit credits

Aggressively cheap for low-latency streaming. Pro is only $5/month for 100,000 credits — but Pro is also the first tier that grants a commercial license and instant voice cloning, so the free tier is strictly an evaluation sandbox. Startup is $49/month for 1.25 million credits; Scale is $299/month for 8 million.

Good: the cheapest paid entry on this list at $5; cloning included from that tier; genuinely low latency.

Not so good: the free tier cannot be published from; a younger company than the clouds; smaller catalogue.

Bottom line: the best price-per-stream here if you are building something latency-sensitive and can live with a smaller vendor.

Microsoft Azure AI Speech — best for Microsoft-stack enterprises

Free 500k neural chars/mo, recurring, heavily throttled · Paid rate varies by region · Unit characters

Azure publishes a free F0 tier of 500,000 neural TTS characters per month, ongoing, alongside 5 audio hours of speech-to-text. It is throttled and not meant for production. We are not quoting Azure's paid rates because the pricing table loads them client-side; read them off the portal for your own region, since Azure rates vary by geography in a way most of this list does not.

Good: deep language and locale coverage; the compliance and residency story large enterprises need; custom neural voice for approved customers.

Not so good: region-dependent pricing you have to look up yourself; the free tier's throttling makes it a demo; the console is a lot.

Bottom line: if your company already runs on Azure, this decision was made for you.

ElevenLabs — best agent tool surface for audio

Covered above on narration quality, but it belongs here too, and honestly so: ElevenLabs ships an official MCP server with a substantially larger tool surface than ours — speech, transcription, speech-to-speech, sound effects, music, voice design, and outbound-calling voice agents — across Claude Desktop, Cursor, Windsurf and OpenAI Agents.

Bottom line: if your agent's job is audio primitives and you want the widest set of them behind one MCP connection, this is the one to install. Our answer below is narrower on audio primitives and wider on finished deliverables.

ListenHub — best for one key across audio and video

Free monthly allowance plus a daily top-up, API key included · Paid entry $12/mo, or $9/mo annually · Unit credits

Our own API is the reason this section got longer, so here is the honest shape of it. It is not a cheaper per-character synthesis endpoint — it is a different unit of work. One POST starts a job that comes back as a finished artefact: a two-host episode with a written script, narration with SRT subtitles, a rendered explainer video, a slide deck, a music bed, an image. Jobs are asynchronous — you create a task and poll it — because a ten-minute render is not an HTTP round trip.

One API key reaches all of it: podcasts, text-to-speech, voice cloning, music, image generation, explainer video, slides, AI video, and content extraction from a URL or file. No second vendor for the video half, no glue code to stitch audio to visuals, one credit balance, one bill. The full API surface is described further down.

Good: finished deliverables rather than audio bytes; audio and video behind one key; four first-party ways in (REST, MCP, Agent Skills, SDK/CLI) so an agent can drive it with no integration written; any account can create a key, so evaluating costs nothing; the async task model suits long renders; the SDK unwraps the response envelope and retries rate limits for you.

Not so good: credits are a coarser unit than characters, so cost-per-request is less predictable than Polly's; the free allowance is small next to Polly's 5 million characters a month, so sustained volume means a subscription or credit packs; ElevenLabs exposes more audio tools over MCP than we do; we publish no latency SLA, and you should not build a real-time voice agent on this.

Bottom line: the right call when you want an agent or a backend to emit finished, publishable content. The wrong call when you want cheap bulk synthesis or sub-second streaming.

A finished program, not a clip

A different product category that gets filed under "text to speech" because people search for it that way. You do not have a script. You have a document, and you want two people discussing it.

NotebookLM — best for free personal listening

Free 3 audio generations/day with any Google account · Paid via Google AI subscription · Unit generations

Free and remarkable at this. You upload sources and it produces a two-host audio overview; Google's own support docs put the free plan at 3 audio generations a day alongside 50 chat queries. If your need is occasional and personal — turning a dense PDF into something for a walk — it is hard to argue with free.

Good: genuinely free; the two-host conversation is convincing; zero setup.

Not so good: you cannot edit the script before it speaks, choose the voices, clone a voice, or remove the branding; output is a single fixed format; not built for publishing on a schedule.

Bottom line: the best free way to listen to your own documents, and not a production tool.

ListenHub — best for a publishable multi-voice program

Free monthly allowance plus a daily top-up · Paid entry $12/mo, or $9/mo annually · Unit credits

This is what we build, so read it with the appropriate suspicion — and read the next section, which is the honest long version including the parts we are bad at. The one-line version: NotebookLM gives you an audio overview, and ListenHub gives you a program you can edit, brand and ship.

Good: the script is editable before anything is spoken; twelve output languages for text-to-speech and voice cloning; the written-to-spoken rewrite that fixes the "read literally" failure above; audio, SRT, video and slides from one credit balance.

Not so good: the AI podcast format itself still generates in English, Chinese and Japanese only; script editing, file exports, cloning and branding removal all need a paid plan, so the free tier is for listening rather than publishing; our per-unit cost cannot touch $4 per million characters; there is no video editing timeline; if you want one voice reading one long script, ElevenLabs does it better.

Bottom line: buy us when the deliverable is a finished multi-voice program built from a source document. For most other jobs on this page, something else on this page is the better answer.

What ListenHub actually does

You brought a report, a paper, a URL or a script. You want something people will actually listen to. Four things happen here that a plain text-to-speech box does not do, and they are the whole reason to pick us.

It rewrites written text into spoken text first. This is the written-to-spoken step, and it addresses the single biggest failure mode in this article. "Fig. 3 shows a 12.4% YoY increase (see Appendix B)" becomes something a human would actually say out loud. Bullet fragments get verbs. Abbreviations get expanded the way a reader would expand them. You can also switch it off and get word-for-word reading when the text is already conversational — a legal disclaimer should be read exactly as written.

It writes dialogue, not two stitched monologues. The turn-taking problem from earlier is a script problem before it is a synthesis problem. Two hosts get a script written as a conversation, with the second speaker responding to what the first one actually said.

You can edit the script before anything is spoken. This is the difference between a toy and a tool. You see the script, you fix the sentence that misreads your product's positioning, you cut the tangent, and only then do you spend credits on audio. Fixing text is free; re-rendering audio is not. Script editing is a subscription feature, from Basic up.

One balance covers the whole output. Audio with SRT subtitles, an explainer video, slides, images — the same credits, no second subscription. File exports and ListenHub branding removal also start at Basic, so the free tier is the place to hear what the tool sounds like rather than the place to publish from.

An API and an agent surface

This is the half of the product that does not show up in a feature grid, and for developers it is the most interesting part. Everything the web app can do is reachable programmatically, through four first-party surfaces that all authenticate with the same API key and draw down the same credit balance.

The REST API. Create a key at Settings → API keys — the format is lh_sk_… — and call https://api.marswave.ai/openapi. Responses use one envelope, { code, message, data }, with a non-zero code for errors. Generation is asynchronous by design: you create a task and poll it, because a full episode or a rendered video takes minutes, not milliseconds. The endpoint set covers podcast generation, text-to-speech, ListenHub Voice for end-to-end audio, music, image generation, explainer video, slides, AI video, speaker lookup, content extraction from a URL or file, and subscription status. The API documentation has the full reference; per-endpoint parameters and language support are specified there rather than here, since that surface moves faster than a blog post.

An MCP server. ListenHub ships a Model Context Protocol server, so an MCP client — Claude Desktop, Cursor, Windsurf, VS Code, Zed — can generate podcasts and narration, browse the voice catalogue and check your account by calling tools instead of hitting HTTP. Each tool wraps a public API endpoint, so the two surfaces stay in step. It runs over stdio by default, or HTTP/SSE when you need it remote, and authenticates from LISTENHUB_API_KEY. Worth knowing: create_podcast polls to completion on your behalf, so one tool call returns a finished episode rather than a task id your agent has to babysit. See the MCP docs.

Agent Skills. The surface we would point a coding agent at first:

npx skills add marswaveai/skills

That installs ListenHub Skills into a project for any tool with Agent Skills support — Claude Code, Cursor, Windsurf, OpenCode. From there the instruction is plain language: turn this article into an explainer video, make a two-host podcast from these three URLs. The agent picks the voices, drives generation, polls for completion and hands back the artefact with a link. There are individual skills for voice, podcast, text-to-speech, explainer video, slides, images, music, transcription and content parsing, so an agent can compose a pipeline instead of calling one monolithic endpoint.

SDK and CLI. @marswave/listenhub-sdk is the official typed JavaScript/TypeScript client — ESM-only, ships its own types, one runtime dependency, Node 20 or newer. It unwraps the response envelope and retries 429s so you do not hand-roll either. It exports two clients, which is a genuinely useful distinction: OpenAPIClient authenticates with an API key for servers, scripts and CI and acts as the key owner, while ListenHubClient uses OAuth user tokens so each request runs as an individual signed-in user — the right choice if you are building something user-facing. @marswave/listenhub-cli wraps the same SDK as a listenhub binary with the same split: listenhub … after an OAuth browser login for interactive work, listenhub openapi … with a key for CI. Docs for the SDK and the CLI.

Picking voices programmatically. Speaker lookup returns more than an ID and a name — each voice carries pitch, speed, traits, styles, suggested scenes, accent and a description, plus localized descriptions. An agent can therefore choose a voice from a brief such as "warm female narrator for an audiobook" rather than hard-coding an ID a human picked once. There is also a plain-text catalogue at /voices.txt for agents that would rather read a file than call an endpoint.

Two things to be clear about. No plan is required — any account can create a key, so you can evaluate the API on the free allowance without a card, and a weekend prototype costs nothing. What you do need is credits: API calls spend the same balance as the web app, so your automated usage and your team's manual usage come out of one pot, and sustained volume means a subscription or a credit pack. The unit is credits rather than characters, which is a coarser thing to forecast than a per-character rate.

Languages and voices

Text-to-speech and voice cloning both run in twelve languages: English, Chinese Mandarin, Japanese, Spanish, Portuguese, French, German, Turkish, Korean, Italian, Thai and Vietnamese. The public voice catalogue carries tags and descriptions per voice so you can filter by language, gender and delivery rather than auditioning names one at a time. There is also a plain-text catalogue at /voices.txt if you are pointing an agent at it.

Be precise about what that does and does not cover, because we have seen this oversold elsewhere: the twelve languages apply to text-to-speech and to cloning. The AI podcast format and the other creation tools still generate in English, Chinese and Japanese. If you need a two-host Spanish podcast today, we are not there yet.

Voice cloning

Clone your own voice from a sample and use it anywhere a catalogue voice works, in any of the twelve languages. Every paid plan includes cloning — Basic 1 saved clone, Pro 4, Max 20 — and saves beyond your plan's allowance cost 300 credits each. The voice cloning guide walks through what makes a good sample; the short version is three minutes of clean speech beats thirty minutes of noisy.

What it costs

Basic is $12/month, or $9/month billed annually, for 1,300 credits and one voice clone. Pro is $24/month, or $19/month annually, for 2,700 credits and four clones. Max is $240/month, or $200/month annually, for 30,000 credits and twenty clones. The API, MCP server and Agent Skills are not tied to a plan — they work on any account, including the free tier, and draw from whatever credits you have.

To translate credits into work: a 5-minute podcast runs about 24 credits, 10 minutes of text-to-speech about 40, a 10-page deck about 150. So Pro is roughly 11 hours of speech a month if you spent it on nothing else. There is a free tier with a monthly allowance plus a small daily top-up; the pricing page carries the current figure, which moves often enough that we would rather point you at it than freeze it here. The credits guide explains how the three credit types stack and the order they are spent in.

Where we are weaker than the tools above

Stated plainly, because you will find this out anyway:

  • Per-unit cost at volume. Polly and Google are $4 per million characters. A subscription cannot compete with that, and we are not going to pretend otherwise.
  • Single-voice long-form polish. ElevenLabs holds prosody better across a very long single-narrator script.
  • Audio tool breadth for agents. ElevenLabs' MCP server exposes more audio primitives than ours — transcription, speech-to-speech, audio isolation, voice design, outbound calling. Ours is aimed at finished deliverables instead.
  • No per-character API metering. You spend credits, not characters, so cost-per-request is harder to forecast than a published $4-per-million rate — and our free allowance is a fraction of Polly's 5 million characters a month.
  • No latency commitment. We publish no SLA and the task model is built for renders, not for real-time turns. Do not put us in a live voice agent.
  • Video editing. Speechify gives you a timeline. We give you an export.
  • Podcast languages. Twelve for TTS and cloning, three for the podcast format.
  • Enterprise procurement. No residency guarantees or custom contracts of the kind Azure and AWS sell.

Free tiers, compared

Free tiers differ in whether they are a product or a demo. The distinction that matters is not size — it is whether you may publish what you make.

ToolFree allowanceRecurring?Publish from free?
Amazon Polly5M chars/mo StandardYes, no cliffYes
Google Cloud4M chars/mo Standard, 1M Chirp 3 HDYesYes
Azure500k neural chars/moYes, throttledYes
Deepgram$200 creditNo, one timeYes
NotebookLM3 generations/dayYesPersonal use
ElevenLabs10k credits/moYesNo
Cartesia20k credits/moYesNo
Speechify600 creditsYesNo
WellSaid3 download min/moYesNo
ListenHubMonthly plus daily top-upYesNo, exports need a plan

The bottom five are perfectly honest about it, and perfectly useless if you intend to publish. We have put ourselves in that group rather than above it, because a free tier you cannot export from is an evaluation sandbox whoever built it. Budget for the paid entry tier from day one — $6, $5, $19, $19 and $12 respectively. Polly's better voice classes are also time-limited: Neural 1 million a month, Long-Form 500,000, Generative 100,000, each for the first 12 months only.

If "free" is your hard requirement and you intend to publish, the shortlist is short: Polly and Google Cloud, and you will be writing code.

Languages beyond English

The cloud platforms win on breadth outright. Google, Azure and Polly each cover dozens of languages with native-speaker voices and locale-specific variants, having built this long before the current wave. ElevenLabs is strong at multilingual output from a single voice — the same cloned voice speaking another language, which is a different and useful thing.

Where the marketing gets slippery is the gap between supported and good. Many vendors quote a large language count where most of the list is served by one multilingual model with a handful of voices tuned for English. OpenAI is candid about this in its own docs. Ask for a sample in your target language, spoken by a voice actually marketed for that language, and have a native speaker listen.

By that standard, here is our own position without the rounding up: twelve languages for text-to-speech and voice cloning, each with voices from the catalogue rather than an English voice with an accent setting — and three languages for the podcast format. If you need Hindi, Arabic or Indonesian narration, we do not serve it, and Google or Azure does.

Which one should you pick?

You are a solo creator publishing narrated video. ElevenLabs Starter at $6/month. Add Speechify only if you want the editing timeline too.

You run an L&D or training team. WellSaid Labs. Minutes are a budgetable unit and the voices survive legal review.

You are an engineer who needs synthesis inside a product. Google Cloud or Polly. Prototype inside the recurring free tier, then pick a voice class deliberately — the difference between $4 and $160 per million characters is a voice-class dropdown.

You are building a voice agent. Deepgram or Cartesia. Latency and cost per stream are the whole game; catalogue size is not.

You want an AI agent to produce finished content on its own. ListenHub Skills or our MCP server. One npx skills add marswaveai/skills and a coding agent can turn a URL into a published episode or an explainer video without an integration being written first. No plan needed to start — a free account gets a key, and you pay in credits once you are past evaluating. If your agent needs audio primitives rather than finished pieces — transcription, sound effects, voice design — install ElevenLabs' MCP server instead, or both.

You are automating a content pipeline in CI. Our SDK or CLI with an API key, or the clouds if all you need is synthesis. listenhub openapi … exists precisely for the scripted case.

You are a student or researcher with dense PDFs. NotebookLM, free. If you start wanting the script editable and the output publishable, that is the point at which you outgrow it.

You publish a show or a content programme from documents. ListenHub. Editable script, multi-voice output, exports without branding, one balance across audio, video and slides.

Your requirement is genuinely $0 and you can code. Polly, then Google Cloud. Everyone else's free tier either forbids publishing or runs out.

Questions people actually ask

What is the best free text-to-speech tool? Amazon Polly, if you can write code: 5 million characters a month, permanently, and you may publish the result. If you cannot write code, NotebookLM is free for personal listening. Most consumer free tiers — ElevenLabs, Cartesia, Speechify, WellSaid and ours — either forbid commercial use or hold exports behind a paid plan.

Can I use AI voices on YouTube or in client work? Only on a tier that grants a commercial licence, and that is usually not the free one. ElevenLabs grants it from $6, Cartesia from $5, Speechify from $19. Check the licence on the exact tier you intend to buy, not the one above it. Platform rules are separate from the vendor's licence, and OpenAI's policies additionally require you to disclose that a voice is AI-generated.

What is voice cloning, and do I need permission? You provide a sample of a voice and the model reproduces its timbre, so it can read any text. You need the consent of the person whose voice it is — your own counts. Cloning someone else's voice without permission is the fastest way to a legal problem, and every reputable vendor's terms prohibit it.

Which tool has the most natural voices? For one voice reading a long script, ElevenLabs, in our experience. But "natural" fails on prosody, literal glyph reading and mangled names far more often than on timbre, so run your own worst paragraph rather than trusting anyone's ranking.

How much should I expect to pay? Roughly $5–$25/month for a creator subscription, or $4–$30 per million characters on an API. Free tiers cover evaluation and personal use; publishing usually starts at the first paid tier.

Can AI-generated speech be detected? Not reliably, and you should not plan around that. Some platforms require synthetic-media disclosure, some vendors require it in their terms, and the reputational cost of being caught not disclosing is worse than the cost of disclosing.

What is the difference between text-to-speech and an AI podcast tool? Text-to-speech takes a script and reads it. An AI podcast tool takes a source document, writes a script — usually as dialogue — and then reads it. If you already have the words, buy text-to-speech. If you have a report and want a conversation about it, that is a different product, and it is the one we build.

Do I need an API or an app? An API if speech is a feature inside your product, or if your volume makes per-character pricing cheaper than a subscription. An app if speech is the deliverable and you would rather not maintain code. Buying an API to make forty episodes by hand is a common and expensive mistake.

Which of these can an AI agent use directly? ElevenLabs and ListenHub both ship official MCP servers, so an agent in Claude Desktop, Cursor or Windsurf can call them as tools with no integration written. ListenHub additionally publishes Agent Skills — npx skills add marswaveai/skills — for Claude Code, Cursor, Windsurf and OpenCode. Everything else on this list is a REST API you or your agent would have to wrap first. The practical difference is what a call returns: ElevenLabs hands your agent audio primitives, ListenHub hands it a finished episode, video or deck.

What does the ListenHub API cost, and is there a free key? Any account can create a key, including a free one — no plan required. What you spend is credits, the same balance the web app draws on, so evaluation is free and sustained volume means a subscription from $12/month or a credit pack. The honest caveat is the unit: credits are coarser than characters, so if you need to forecast cost per request to the cent, a per-character vendor like Polly or Google Cloud is easier to model.

Can I call ListenHub on behalf of my own signed-in users? Yes, and this is what the two SDK clients are for. OpenAPIClient uses an API key and acts as the key owner — right for servers, scripts and CI. ListenHubClient uses OAuth user tokens so each request runs under an individual user's account, which is what you want in a user-facing app. Never ship an API key in browser or mobile code.

Which languages does ListenHub support? Twelve for text-to-speech and voice cloning: English, Chinese Mandarin, Japanese, Spanish, Portuguese, French, German, Turkish, Korean, Italian, Thai, Vietnamese. Three — English, Chinese, Japanese — for the AI podcast format. For the API and MCP surfaces, check the reference docs per endpoint rather than assuming the web app's range.

When we would send you elsewhere

Six cases, stated plainly:

You need one voice reading a long script beautifully. Buy ElevenLabs. Long-form prosody is their specialty and $6/month is not a real barrier.

You are an engineer with volume and a budget. Go to Google Cloud or Polly. Nobody selling a subscription can beat $4 per million characters, and the recurring free tiers mean your prototype costs nothing.

Your agent needs audio primitives, not finished pieces. Install ElevenLabs' MCP server. Transcription, speech-to-speech, sound effects, voice design and outbound calling are all there, and our tool surface is deliberately narrower.

You are building anything real-time. Deepgram or Cartesia. We publish no latency SLA and our task model is built for renders.

You need a language we do not serve, or a two-host podcast outside English, Chinese and Japanese. Google or Azure for narration breadth. Our podcast format has not caught up with our voice catalogue yet.

You want a document turned into a listenable summary for yourself. Use NotebookLM. It is free and good, and you do not need script editing, branding removal or voice cloning for a private commute listen.

We are the better answer in two situations: when you want a finished, publishable multi-voice program out of a source document with the script editable before anything is spoken, and when you want an agent or a backend to produce that kind of output — audio and video alike — behind a single key.

Run your own test before you commit

Twenty minutes will tell you more than any comparison article, including this one.

  1. Take your worst real paragraph — product names, acronyms, numbers, a parenthetical aside. Not the vendor's demo text.
  2. Run it through the three tools on your shortlist, unedited.
  3. Listen at 1x on headphones, all the way through. Speed-listening hides pacing problems.
  4. Count the failures from earlier: wrong stress, glyphs read literally, mangled names, dead turn-taking, uniform gaps.
  5. Only then look at price. A tool that mispronounces your company name is not cheap at any rate.
  6. Check the commercial license on the exact tier you intend to buy, not the one above it.

If you want to try ours in that test, text-to-speech and AI podcast both run on the free tier without a card. If you would rather evaluate from a terminal than a browser, npx skills add marswaveai/skills puts it in front of your coding agent, and the API docs are open to read before you pay for anything.

And if a competitor wins on your text, they have earned it.

← Back to blog