Every generation model ListenHub runs. Pick one to open its tool, or take the model ID and call it from the API.
Text, image and first/last-frame input, plus a reference video to drive camera work. The steadiest multi-subject consistency we have; first choice for 1080p.
The low-latency lane of the same model — drafts, batches and quick shot tests. No 1080p.
Generates from first and last frames rather than reference images. Lip sync and motion transfer both run on it.
The standard Wan 3.0 lane. Duration and resolution ceilings match Prime; the difference is end-to-end speed.
Steady human motion and camera movement, with native audio output. Does not accept video input.
Rich camera language; suits narrative sequences and shots with obvious movement.