Explainer & Slides
One endpoint family that turns text or a URL into a narrated page deck — explainer, story, or presentation slides — and renders it to video.
Explainer videos, story videos, and presentation slides are three modes of one endpoint family. You send a source and a voice, ListenHub writes the narration and generates a visual per page, and you either use the pages as raw material or render them into a narrated video.
The three endpoints are shared by every mode:
| Endpoint | Purpose |
|---|---|
POST /v1/storybook/episodes | Create an episode. mode selects explainer, story, or slides |
GET /v1/storybook/episodes/{episodeId} | Poll progress, then read the pages and asset URLs |
POST /v1/storybook/episodes/{episodeId}/video | Render the finished pages into a video |
All endpoints on this page use the OpenAPI base URL https://api.marswave.ai/openapi and
authenticate with your API key via the Authorization: Bearer $LISTENHUB_API_KEY header. Create
keys at listenhub.ai/settings/api-keys.
Choose a mode
mode is optional and defaults to info. It is the only field that differs between the three
modes — everything else on this page applies to all of them.
mode | Produces | Page 1 | Best for |
|---|---|---|---|
info (default) | Infographics, illustrations, data visualizations | Magazine-style cover | Knowledge explainers, product intros |
story | Story scene illustrations | Story cover | Story sharing, case studies |
slides | PPT layouts (grid, process flow, big-number hero) | Presentation title page | Meeting presentations, business reports, conference talks |
Explainer & Story
The info and story modes: what each one draws, and when to pick which.
Slides
The slides mode: presentation layouts, and how to get a PPT file out.
Workflow
Create the episode
Call POST /v1/storybook/episodes with your source, a speaker, and the mode you want. Save the
returned episodeId.
Poll for completion
Poll GET /v1/storybook/episodes/{episodeId} every 10 seconds, after an initial 60-second wait,
until processStatus is success.
Use the raw materials (optional)
pages[] holds the generated images and narration scripts. If that is all you need, stop here.
Render the video
Call POST /v1/storybook/episodes/{episodeId}/video to combine the pages into a narrated video.
Poll the video
Poll until videoStatus is success, then download videoUrl.
Create an episode
POST /v1/storybook/episodes
curl -X POST "https://api.marswave.ai/openapi/v1/storybook/episodes" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"sources": [
{"type": "url", "content": "https://example.com/article"}
],
"speakers": [
{"speakerId": "<SPEAKER_ID>"}
],
"language": "en",
"mode": "info"
}'const response = await fetch(
'https://api.marswave.ai/openapi/v1/storybook/episodes',
{
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.LISTENHUB_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
sources: [{ type: 'url', content: 'https://example.com/article' }],
speakers: [{ speakerId: '<SPEAKER_ID>' }],
language: 'en',
mode: 'info',
}),
},
)
const { data } = await response.json()
const episodeId = data.episodeIdimport os
import requests
response = requests.post(
'https://api.marswave.ai/openapi/v1/storybook/episodes',
headers={'Authorization': f'Bearer {os.environ["LISTENHUB_API_KEY"]}'},
json={
'sources': [{'type': 'url', 'content': 'https://example.com/article'}],
'speakers': [{'speakerId': '<SPEAKER_ID>'}],
'language': 'en',
'mode': 'info',
},
)
episode_id = response.json()['data']['episodeId']Response:
{
"code": 0,
"message": "",
"data": { "episodeId": "665f1d4e8b3a3f001234abcd" }
}Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
sources | array(1) | Yes | Content source. Exactly 1 item |
sources[].type | string | Yes | "text" or "url" |
sources[].content | string | Yes | Text content, or the URL itself when type is "url" |
sources[].uri | string | No | Accepted but ignored — the server derives uri from content for url sources |
sources[].metadata | object | No | Source metadata |
speakers | array(1) | Yes¹ | Voice config. At most 1 item |
speakers[].speakerId | string | Yes | Speaker ID (see Speakers) |
skipAudio | boolean | No | Defaults to false. When true, only images and text are produced — no narration audio |
language | string | No | Language code, e.g. "en", "zh". Defaults to en — it is not inferred from the source |
mode | string | No | "info" (default), "story", or "slides" |
style | string | No | Visual style ID. Omit it to use the mode's default style; the accepted IDs are not part of the public contract |
¹ speakers is required unless skipAudio is true, in which case it may be omitted.
language is not detected from your content. Omitting it produces an English episode whatever the
source language is, so send it explicitly for any non-English source.
Set skipAudio: true when you only want the visuals and the script — the page images and
narration text still come back on the episode, without the audio render.
Poll the episode
GET /v1/storybook/episodes/{episodeId}
Poll with the returned episodeId until processStatus is success.
curl "https://api.marswave.ai/openapi/v1/storybook/episodes/{episodeId}" \
-H "Authorization: Bearer $LISTENHUB_API_KEY"Response (when processStatus is success):
{
"code": 0,
"message": "",
"data": {
"episodeId": "{episodeId}",
"createdAt": 1700000000,
"mode": "info",
"processStatus": "success",
"credits": 30,
"title": "How AI Is Changing the World",
"cover": "https://assets.listenhub.app/covers/{episodeId}.png",
"audioUrl": "https://assets.listenhub.app/storybook/{episodeId}.mp3",
"audioDuration": 180,
"videoUrl": "",
"videoStatus": "not_generated",
"pages": [
{
"text": "Artificial intelligence has transformed industries worldwide...",
"pageNumber": 1,
"imageUrl": "https://assets.listenhub.app/pages/{episodeId}-1.png",
"audioTimestamp": 0
},
{
"text": "From healthcare to finance, AI applications continue to expand...",
"pageNumber": 2,
"imageUrl": "https://assets.listenhub.app/pages/{episodeId}-2.png",
"audioTimestamp": 25.3
}
]
}
}Raw materials: each item in pages[] carries an imageUrl (the generated visual), text (the
narration script), and audioTimestamp (where that page starts in audioUrl). Download them and
build or edit your own deck without ever rendering a video.
Response fields
| Field | Type | Description |
|---|---|---|
episodeId | string | The episode identifier you polled with |
mode | string | The mode this episode was created in |
processStatus | string | See below |
videoStatus | string | See below |
credits | number | Credits this episode has consumed so far |
failCode | number | Failure code, present when processStatus is fail |
message | string | Human-readable detail about the current status |
title / cover | string | Generated title and cover image |
audioUrl / audioDuration | string / number | Narration audio and its length in seconds |
videoUrl | string | Rendered video, once videoStatus is success |
pages[] | array | text, pageNumber, imageUrl, audioTimestamp per page |
Credits are reported per episode rather than quoted up front — see Credits & pricing for how the platform bills generation.
processStatus
| Value | Meaning |
|---|---|
pending | Processing |
success | Complete |
fail | Failed — read failCode and message, and see Error Handling |
videoStatus
| Value | Meaning |
|---|---|
not_generated | Video not yet triggered |
pending | Video generating |
success | Video ready (videoUrl available) |
fail | Video generation failed |
Generation typically takes 2–5 minutes. Recommended polling: wait 60 seconds, then poll every 10 seconds.
Render the video
POST /v1/storybook/episodes/{episodeId}/video
Trigger video generation for a completed episode. processStatus must be success first.
curl -X POST "https://api.marswave.ai/openapi/v1/storybook/episodes/{episodeId}/video" \
-H "Authorization: Bearer $LISTENHUB_API_KEY"Response:
{
"code": 0,
"message": "",
"data": { "success": true }
}After triggering, poll GET /v1/storybook/episodes/{episodeId} until videoStatus is success and
read videoUrl.