ホーム
アセット
マイ作品
Models
ボイスライブラリ
開発者
ブログ
日本語
料金
ログイン

Models

Every generation model ListenHub runs. Pick one to open its tool, or take the model ID and call it from the API.

API docs
Best for videoSeedance 2.0 ProThe only one with both reference video and native audio
Best for images
Nano Banana Pro
Steady text rendering and multi-reference consistency
Best for voiceLFlowSpeechOne engine behind podcasts, voiceover and dialogue

Capability: reference video · first/last frame · native audio · multi-reference · precise edit · multi-speaker · emotion · streaming
Access: has a tool page · API only

ByteDance: Seedance 2.0 ProPopular

Text, image and first/last-frame input, plus a reference video to drive camera work. The steadiest multi-subject consistency we have; first choice for 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
ByteDance: Seedance 2.0 Fast

The low-latency lane of the same model — drafts, batches and quick shot tests. No 1080p.

6 ratiosInput: text · image · videoBilling: 按分辨率 × 时长,生成前实时估价Tools: ai-video
PPixVerse

Generates from first and last frames rather than reference images. Lip sync and motion transfer both run on it.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: lip-sync
Alibaba: Wan 3.0

The standard Wan 3.0 lane. Duration and resolution ceilings match Prime; the difference is end-to-end speed.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
Alibaba: Wan 3.0 Prime

Steady human motion and camera movement, with native audio output. Does not accept video input.

5 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video
MiniMax: MiniMax H3

Rich camera language; suits narrative sequences and shots with obvious movement.

6 ratiosInput: text · imageBilling: 按秒计费,分辨率越高越贵,生成前实时估价Tools: ai-video