Every generation model ListenHub runs. Pick one to open its tool, or take the model ID and call it from the API.
Text, image and first/last-frame input, plus a reference video to drive camera work. The steadiest multi-subject consistency we have; first choice for 1080p.
The speech engine every voice tool shares. Standard voices run on our own engine, Pro voices on ElevenLabs; voices are a layer under it.
Reliable quality and text rendering, with good consistency across multiple reference images. The everyday default.
The fast lane of the same family, with the widest aspect-ratio support including extremes like 1:4 and 8:1.
Strong instruction following — it understands edits and local replacements. Up to 4 reference images.
The official channel, with an extra quality setting. High quality and 4K need a subscription.
Good grasp of Chinese-language scenes and layout. Up to 9 reference images.
The standard lane of the same family — 1K and 2K output at a lower cost.
Built for precise edits, with up to 10 reference images. Ratios are derived from pixel size, so framing is barely limited.
The low-latency lane of the same model — drafts, batches and quick shot tests. No 1080p.
Generates from first and last frames rather than reference images. Lip sync and motion transfer both run on it.
The standard Wan 3.0 lane. Duration and resolution ceilings match Prime; the difference is end-to-end speed.
Steady human motion and camera movement, with native audio output. Does not accept video input.
Rich camera language; suits narrative sequences and shots with obvious movement.
A voice-acting model billed per second. Its tool page is called AI Voice.
Clone your own voice from one recording; the result shows up under My voices in FlowSpeech.
Sings lyrics you supply, with remix, extension and stem separation around it — for when the track exists and needs more work.
A style description is enough for a whole track; lyrics are optional, and one flag makes it instrumental.