Home
Assets
My Works
Models
Voice Library
Developer
Blog
English
Pricing
Sign In

Models

Every generation model ListenHub runs. Pick one to open its tool, or take the model ID and call it from the API.

API docs
Best for videoSeedance 2.0 ProThe only one with both reference video and native audioBest for imagesNano Banana ProSteady text rendering and multi-reference consistencyBest for voiceLFlowSpeechOne engine behind podcasts, voiceover and dialogue

Capability: reference video · first/last frame · native audio · multi-reference · precise edit · multi-speaker · emotion · streaming
Access: has a tool page · API only

ByteDance: Seedance 2.0 ProPopular

Text, image and first/last-frame input, plus a reference video to drive camera work. The steadiest multi-subject consistency we have; first choice for 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
ByteDance: Seedance 2.0 Fast

The low-latency lane of the same model — drafts, batches and quick shot tests. No 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
PPixVerse

Generates from first and last frames rather than reference images. Lip sync and motion transfer both run on it.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: lip-sync
Alibaba: Wan 3.0

The standard Wan 3.0 lane. Duration and resolution ceilings match Prime; the difference is end-to-end speed.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
Alibaba: Wan 3.0 Prime

Steady human motion and camera movement, with native audio output. Does not accept video input.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
MiniMax: MiniMax H3

Rich camera language; suits narrative sequences and shots with obvious movement.

6 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video