Home
Assets
My Works
Models
Voice Library
Developer
Blog
English
Pricing
Sign In

Models

Every generation model ListenHub runs. Pick one to open its tool, or take the model ID and call it from the API.

API docs
Best for videoSeedance 2.0 ProThe only one with both reference video and native audioBest for imagesNano Banana ProSteady text rendering and multi-reference consistencyBest for voiceLFlowSpeechOne engine behind podcasts, voiceover and dialogue

Capability: reference video · first/last frame · native audio · multi-reference · precise edit · multi-speaker · emotion · streaming
Access: has a tool page · API only

ByteDance: Seedance 2.0 ProPopular

Text, image and first/last-frame input, plus a reference video to drive camera work. The steadiest multi-subject consistency we have; first choice for 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
LListenHub: FlowSpeechPopular

The speech engine every voice tool shares. Standard voices run on our own engine, Pro voices on ElevenLabs; voices are a layer under it.

Input: textBilling: 按音频时长,约 4 积分 / 分钟Tools: text-to-speech · podcast · multi-speaker-tts
Google: Nano Banana Pro

Reliable quality and text rendering, with good consistency across multiple reference images. The everyday default.

1K / 2K / 4K · 10 ratiosInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
Google: Nano Banana Flash

The fast lane of the same family, with the widest aspect-ratio support including extremes like 1:4 and 8:1.

1K / 2K / 4K · 14 ratiosInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
OpenAI: GPT-Image-2

Strong instruction following — it understands edits and local replacements. Up to 4 reference images.

1K / 2K / 4K · 10 ratios · up to 4 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
OpenAI: GPT-Image-2.5 Flare
1K / 2K / 4K · 10 ratios · up to 4 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化API docs
OpenAI: GPT-Image-2.5 Sunburst
1K / 2K / 4K · 10 ratios · up to 4 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化API docs
OpenAI: GPT-Image-2 Pro

The official channel, with an extra quality setting. High quality and 4K need a subscription.

1K / 2K / 4K · 10 ratios · up to 4 reference imagesInput: text · imageBilling: 按张计费,随尺寸档与质量变化Tools: ai-image
Alibaba: Wan 2.7 Image Pro

Good grasp of Chinese-language scenes and layout. Up to 9 reference images.

1K / 2K / 4K · 5 ratios · up to 9 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
Alibaba: Wan 2.7 Image

The standard lane of the same family — 1K and 2K output at a lower cost.

1K / 2K · 5 ratios · up to 9 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
ByteDance: Seedream 5.0 Pro

Built for precise edits, with up to 10 reference images. Ratios are derived from pixel size, so framing is barely limited.

1K / 2K · 14 ratios · up to 10 reference imagesInput: text · imageBilling: 按张计费,随尺寸档变化Tools: ai-image
ByteDance: Seedance 2.0 Fast

The low-latency lane of the same model — drafts, batches and quick shot tests. No 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
PPixVerse

Generates from first and last frames rather than reference images. Lip sync and motion transfer both run on it.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: lip-sync
Alibaba: Wan 3.0

The standard Wan 3.0 lane. Duration and resolution ceilings match Prime; the difference is end-to-end speed.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
Alibaba: Wan 3.0 Prime

Steady human motion and camera movement, with native audio output. Does not accept video input.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
MiniMax: MiniMax H3

Rich camera language; suits narrative sequences and shots with obvious movement.

6 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
ByteDance: Seed Audio

A voice-acting model billed per second. Its tool page is called AI Voice.

Input: textBilling: 0.5 积分 / 秒,最低 1 积分Tools: ai-voice
LListenHub: Voice Clone

Clone your own voice from one recording; the result shows up under My voices in FlowSpeech.

Input: text · audioBilling: 见定价说明Tools: voice-cloning
MurekaAPI only

Sings lyrics you supply, with remix, extension and stem separation around it — for when the track exists and needs more work.

Input: textBilling: 见 API 文档API docs
SunoAPI only

A style description is enough for a whole track; lyrics are optional, and one flag makes it instrumental.

Input: textBilling: 见 API 文档API docs