AI Tools · Audio · Video · Image · Seedance 2.0 · Nano Banana
Start with the material you have: a topic, document, link, image, recorded footage or finished script. Speech tools share the ListenHub voice library, while each video page lists the models that support its input type.
Generate video from a script, a photo, or a reference clip.
Explainer VideoAudioScenes and narration, built together from one topic.
Describe an image, then edit just the parts you don’t like
SlidesVideoGenerate a slide deck, with narration if you want.
AI Podcast turns a topic, document or link into a finished episode, with one host or two. Text to Speech reads a script you already have, word for word or lightly rewritten into spoken language first. Multi-Speaker Voiceover gives each role in a dialogue script its own voice, and AI Voice is the page for when tone, emotion and pacing are the point. All of them export an audio file with SRT subtitles; AI Voice exports audio only.
Start at AI Video if you are not sure — it takes a script, a still image or reference footage, and models switch inside one editor. Explainer Video is the pick when the job is to explain something, since the scenes and the narration are built together. Ad Generator, Promo Video, Motion Transfer and Lip Sync each do one job on a fixed model.
AI Image generates from a prompt or up to five reference pictures; when part of the result is wrong you select that region, describe the change, and only that area is redrawn and saved as a new version instead of replacing the old one. Slides builds a deck where every page has real content and a narration script, and exports PPTX and PDF, plus video, audio and SRT subtitles when narration is on.
Record or upload a 25–35 second sample in Voice Cloning and the voice appears in the voice picker for AI Podcast, Text to Speech, Slides and Explainer Video, and can also be selected in AI Voice and Multi-Speaker Voiceover. The video tools do not take a cloned voice.
The tool pages are grouped by output type. Audio: AI Podcast, Text to Speech, Voice Cloning, AI Voice and Multi-Speaker Voiceover. Video: AI Video, Explainer Video, Text to Video, Image to Video, Reference to Video, Ad Generator, Promo Video, Motion Transfer and Lip Sync. Image: AI Image and Slides.
A voice you build in Voice Cloning shows up in the voice picker for AI Podcast, Text to Speech, Slides and Explainer Video, and can also be selected in AI Voice and Multi-Speaker Voiceover. The video tools do not take a cloned voice. The voice library is shared by every tool that speaks, and every voice in it is free to listen to.
There are two tiers. Text to Speech, Multi-Speaker Voiceover, AI Voice and Voice Cloning cover 12 languages: English, Chinese Mandarin, Japanese, Spanish, Portuguese, French, German, Turkish, Korean, Italian, Thai and Vietnamese. AI Podcast, Slides and Explainer Video generate in English, Chinese Mandarin and Japanese.
It depends on what you already have. Text to Video starts from a written prompt alone. Image to Video takes a still as the first frame, and adding a second still as the last frame turns it into a transition. Reference to Video is for when references should decide the person, the product or the style, with the prompt deciding how the shot moves; references and frames stay on separate pages. Those three run up to 15 seconds. Ad Generator defaults to 30 seconds, Promo Video to 20 with 15 and 30 second options, and Motion Transfer and Lip Sync take their length from the clip you upload.
Video runs on Seedance 2.0 Pro, Seedance 2.0 Fast, PixVerse and HappyHorse, switchable in one editor, with each page listing only the models that accept its input and each model offering only the durations and ratios it supports. Seedance 2.5 is not currently listed in the model menu. Ad Generator runs on the PixVerse ad agent, Promo Video on the PixVerse promo agent, and Motion Transfer and Lip Sync on PixVerse. Images use Nano Banana, GPT-Image-2 and Seedream 5.0 Pro in one editor, and AI Voice runs on seed-audio.
The tools that speak export an audio file with SRT subtitles, AI Voice audio only. Slides exports PPTX and PDF, plus video, audio and SRT when narration is on. The video tools download the finished clip, and AI Image downloads the original file. Generations are paid in credits: the cost depends on the model, the duration and the quality, and the exact figure appears next to the generate button before you confirm.