Home
Audio5
Video9
Image2
All Tools
Assets
Voice Library
Developer Docs
Blog
English
Sign In
Download App

AI Tools · Audio · Video · Image · Seedance 2.0 · Nano Banana

AI Tools for Audio, Video and Images

Choose tools by output type

Start with the material you have: a topic, document, link, image, recorded footage or finished script. Speech tools share the ListenHub voice library, while each video page lists the models that support its input type.

  • Podcasts, voiceovers and cloned voices
  • Video from text, an image or reference footage
  • Ads and promos built by the PixVerse agents
  • Images you can fix region by region
  • Narrated decks and explainer videos
AudioVideoImage
Audio5
AI Podcast

Turn a document into a podcast episode, solo or multi-voice.

Text to Speech

Turn text into speech, with many voices and languages

Voice Cloning

Clone your own voice and reuse it across projects

Scenes

AI Voice

Cast a different voice for each speaker in one script.

Multi-Speaker Voiceover

Paste a dialogue script, give every speaker their own voice

Video9
AI Video

Generate video from a script, a photo, or a reference clip.

Explainer VideoAudio

Scenes and narration, built together from one topic.

AI Video · By source

Text to Video

Write one line, get a video

Image to Video

Seedance 2.0 at full strength, portrait first frames supported

Reference to Video

Seedance 2.0 at full strength, portrait references supported

AI Video · More playbooks

Ad Generator

Product photos in, a vertical ad out

Promo Video

Your photos and your footage, cut into one promo

Motion Transfer

Your character, performing that move

Lip SyncAudio

Give the person on screen a new line

Image2
AI Image

Describe an image, then edit just the parts you don’t like

SlidesVideo

Generate a slide deck, with narration if you want.

Where to go instead

  • AI Image to Video Generator
  • AI Reference to Video
  • AI Video Generator

Which tool for which job

Something to listen to

AI Podcast turns a topic, document or link into a finished episode, with one host or two. Text to Speech reads a script you already have, word for word or lightly rewritten into spoken language first. Multi-Speaker Voiceover gives each role in a dialogue script its own voice, and AI Voice is the page for when tone, emotion and pacing are the point. All of them export an audio file with SRT subtitles; AI Voice exports audio only.

Something to watch

Start at AI Video if you are not sure — it takes a script, a still image or reference footage, and models switch inside one editor. Explainer Video is the pick when the job is to explain something, since the scenes and the narration are built together. Ad Generator, Promo Video, Motion Transfer and Lip Sync each do one job on a fixed model.

A picture or a deck

AI Image generates from a prompt or up to five reference pictures; when part of the result is wrong you select that region, describe the change, and only that area is redrawn and saved as a new version instead of replacing the old one. Slides builds a deck where every page has real content and a narration script, and exports PPTX and PDF, plus video, audio and SRT subtitles when narration is on.

In your own voice

Record or upload a 25–35 second sample in Voice Cloning and the voice appears in the voice picker for AI Podcast, Text to Speech, Slides and Explainer Video, and can also be selected in AI Voice and Multi-Speaker Voiceover. The video tools do not take a cloned voice.

AI tools FAQ

Which tools are on ListenHub?

The tool pages are grouped by output type. Audio: AI Podcast, Text to Speech, Voice Cloning, AI Voice and Multi-Speaker Voiceover. Video: AI Video, Explainer Video, Text to Video, Image to Video, Reference to Video, Ad Generator, Promo Video, Motion Transfer and Lip Sync. Image: AI Image and Slides.

Which tools can use a voice I cloned?

A voice you build in Voice Cloning shows up in the voice picker for AI Podcast, Text to Speech, Slides and Explainer Video, and can also be selected in AI Voice and Multi-Speaker Voiceover. The video tools do not take a cloned voice. The voice library is shared by every tool that speaks, and every voice in it is free to listen to.

What languages can the tools output?

There are two tiers. Text to Speech, Multi-Speaker Voiceover, AI Voice and Voice Cloning cover 12 languages: English, Chinese Mandarin, Japanese, Spanish, Portuguese, French, German, Turkish, Korean, Italian, Thai and Vietnamese. AI Podcast, Slides and Explainer Video generate in English, Chinese Mandarin and Japanese.

Which video page should I start on, and how long can a clip be?

It depends on what you already have. Text to Video starts from a written prompt alone. Image to Video takes a still as the first frame, and adding a second still as the last frame turns it into a transition. Reference to Video is for when references should decide the person, the product or the style, with the prompt deciding how the shot moves; references and frames stay on separate pages. Those three run up to 15 seconds. Ad Generator defaults to 30 seconds, Promo Video to 20 with 15 and 30 second options, and Motion Transfer and Lip Sync take their length from the clip you upload.

Which models are behind the tools?

Video runs on Seedance 2.0 Pro, Seedance 2.0 Fast, PixVerse and HappyHorse, switchable in one editor, with each page listing only the models that accept its input and each model offering only the durations and ratios it supports. Seedance 2.5 is not currently listed in the model menu. Ad Generator runs on the PixVerse ad agent, Promo Video on the PixVerse promo agent, and Motion Transfer and Lip Sync on PixVerse. Images use Nano Banana, GPT-Image-2 and Seedream 5.0 Pro in one editor, and AI Voice runs on seed-audio.

What can I export, and what does a generation cost?

The tools that speak export an audio file with SRT subtitles, AI Voice audio only. Slides exports PPTX and PDF, plus video, audio and SRT when narration is on. The video tools download the finished clip, and AI Image downloads the original file. Generations are paid in credits: the cost depends on the model, the duration and the quality, and the exact figure appears next to the generate button before you confirm.