AI Video Tools · Seedance 2.0 · PixVerse · HappyHorse
Choose a workflow based on your source material and intended output. General generation pages let you select compatible models; task-specific pages use the model configured for that workflow.
Generate video from a script, a photo, or a reference clip.
Explainer VideoAudioScenes and narration, built together from one topic.
Text to VideoWrite one line, get a video
Image to VideoSeedance 2.0 at full strength, portrait first frames supported
Reference to VideoSeedance 2.0 at full strength, portrait references supported
Ad GeneratorProduct photos in, a vertical ad out
Promo VideoYour photos and your footage, cut into one promo
Motion TransferYour character, performing that move
Lip SyncAudioGive the person on screen a new line
The head page of the category takes a script, a still image, or reference footage and holds all four models in one editor, in vertical, square, or landscape. It is built for short clips, so a longer story is split into segments and generated one at a time.
Text to Video works from words alone. Image to Video grows the shot out of a still you upload as the first frame, and a last frame turns it into a transition. Reference to Video is the one to reach for when a person, a product, or a style has to stay consistent: the references fix the look, the prompt drives the motion.
Ad Generator turns product photos into a vertical ad, Promo Video cuts the same photos together with footage you shot, Motion Transfer makes the figure in your picture perform a movement from a clip, and Lip Sync matches an on-screen mouth to a voice track you upload. Each of the four runs one fixed model, so there is no model list.
Give it a topic, a document, or a link and the scenes and the narration are built together, in infographic or storytelling mode, in English, Chinese Mandarin, or Japanese. The narration voice comes from the voice library or from a voice you cloned yourself, and the video, the audio track, and SRT subtitles export separately.
Start from what you have. Words only: Text to Video. A still image: Image to Video, where it becomes the first frame. A person, product, or style that has to stay consistent: Reference to Video. Product photos for paid social: Ad Generator. Those photos plus your own footage: Promo Video. A subject picture and a movement to copy: Motion Transfer. A clip that needs new speech: Lip Sync. A subject that needs explaining: Explainer Video. If you are still weighing it up, AI Video accepts a script, a still, or reference footage on one page.
AI Video, Text to Video and Image to Video hold four models in one editor: Seedance 2.0 Pro, Seedance 2.0 Fast, PixVerse and HappyHorse. Reference to Video runs three of them, because PixVerse builds from frames rather than references. The four playbook pages show one fixed model instead of a list: the PixVerse ad agent on Ad Generator, the PixVerse promo agent on Promo Video, and PixVerse on Motion Transfer and Lip Sync. Seedance 2.0 supports portrait references. Seedance 2.5 is not currently listed in the model menu.
Text to Video, Image to Video and Reference to Video top out at 15 seconds. The playbook pages have their own durations: Ad Generator defaults to 30 seconds, Promo Video to 20 seconds with 15 and 30 second options, Motion Transfer follows the motion clip you upload, and Lip Sync inherits length, quality and aspect ratio from the source clip. For a longer story, split it into segments and generate them one at a time.
It depends on the page. Seedance 2.0 Pro and Seedance 2.0 Fast generate sound along with the picture, while HappyHorse generates picture only. On AI Video the generated visuals can be paired with an AI voiceover and on-screen text. The ad agent adds sound effects to the cut and the promo agent handles rhythm and sound. Explainer Video puts narration on every scene. Lip Sync works the other way round: you supply the voice track yourself, as an MP3 or WAV between 5 and 60 seconds.
Image to Video uses your picture as the first frame of the shot, and a second one can be added as the last frame so the motion between them is generated. On Reference to Video the references, up to nine images, a reference clip, or one of each, decide the person, the product and the style, but not the opening frame: that comes from the prompt. References and frames stay on separate pages, so they cannot be mixed in one run.
Video runs are priced in credits, and the amount depends on the model, the duration and the quality you picked. The exact figure appears next to the generate button before you confirm, and on Text to Video the credits are only spent once you confirm.