Image Generation
Generate images from text prompts and reference images across the Nano Banana, GPT-Image, Seedream, and Wan 2.7 model families, synchronously or as async tasks.
Image Generation turns a text prompt — optionally guided by reference images — into one or more
images. Nine models across four families share one request schema and one set of endpoints; you pick
a model with the model field.
You can call it three ways:
- Synchronous —
POST /v1/images/generationblocks until the image is ready and returns the raw model output (base64 image data) in the response body. - Asynchronous —
POST /v1/images/generation/asyncreturns ataskIdimmediately; pollGET /v1/images/generation/tasks/{taskId}for hosted image URLs. - Estimate first —
POST /v1/images/generation/estimate-creditsreturns the credit cost and whether your account can generate, without spending anything.
All endpoints on this page use the OpenAPI base URL https://api.marswave.ai/openapi and
authenticate with your API key via the Authorization: Bearer $LISTENHUB_API_KEY header. Create
keys at listenhub.ai/settings/api-keys.
Choose a model
Use this table to pick a model, then open its family page for the exact limits, quirks, and examples.
| Model | Family | Best for | Sizes | Aspect ratios | Max reference images |
|---|---|---|---|---|---|
gpt-image-2 (default) | GPT-Image | Strong prompt following at a fixed per-size price | 1K, 2K, 4K | The 10 standard ratios | 4 |
gpt-image-2.5-flare | GPT-Image | A second GPT-Image style at the same price and limits | 1K, 2K, 4K | The 10 standard ratios | 4 |
gpt-image-2.5-sunburst | GPT-Image | A third GPT-Image style at the same price and limits | 1K, 2K, 4K | The 10 standard ratios | 4 |
gpt-image-2-official | GPT-Image | The only model that honors quality; priced per token | 1K, 2K, 4K | The 10 standard ratios | 4 |
gemini-3-pro-image | Nano Banana | The most detailed output; Nano Banana Pro | 1K, 2K, 4K | The 10 standard ratios | 14 |
gemini-3.1-flash-image | Nano Banana | Faster, cheaper generation; the only Nano Banana tier that accepts 1:4 / 4:1 / 1:8 / 8:1 | 1K, 2K, 4K | All 14 | 14 |
seedream-5-0-pro | Seedream | Precise region editing from coordinates and color codes | 1K, 2K | All 14 | 10 |
wan2.7-image | Wan 2.7 | Cost-efficient generation on five common ratios | 1K, 2K | 5 ratios | 9 |
wan2.7-image-pro | Wan 2.7 | The same five ratios with 4K output | 1K, 2K, 4K | 5 ratios | 9 |
gpt-image-2 is the default when model is omitted. The 10 standard ratios are 1:1, 2:3,
3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9 — every ratio except the four extreme
ones. See Aspect ratios for the full rules.
Nano Banana
gemini-3-pro-image and gemini-3.1-flash-image: Google's detail-first and speed-first models, plus the four extreme ratios.
GPT-Image
gpt-image-2, the 2.5 Flare and Sunburst variants, and GPT-Image-2 Pro — the only one that honors quality.
Seedream
seedream-5-0-pro: edit a named region of an existing image by writing coordinates into the prompt.
Wan 2.7
wan2.7-image and wan2.7-image-pro: five ratios, strictly validated, with 4K on the Pro tier.
Workflow
Estimate the cost
Call POST /v1/images/generation/estimate-credits with the same body you intend to generate with.
It returns credits and canGenerate without spending anything. Credit cost varies by model and
size — read it from this endpoint rather than hardcoding it.
Generate
Call POST /v1/images/generation to block until the image is ready, or
POST /v1/images/generation/async to get a taskId back immediately.
Collect the result
The synchronous endpoint returns base64 image data in the response body. The async endpoint stores
the image for you — poll GET /v1/images/generation/tasks/{taskId} until status is success and
read the hosted images[].url.
Request parameters
These parameters apply to POST /v1/images/generation, POST /v1/images/generation/async, and
POST /v1/images/generation/estimate-credits — the three endpoints share one request schema.
| Field | Type | Required | Description |
|---|---|---|---|
provider | string | Yes¹ | Model provider: google, openai, bytedance, alibaba, or quotaflow |
model | string | No | Model name. Defaults to gpt-image-2. Legacy preview IDs are normalized to their GA ID. See Choose a model |
prompt | string | Yes² | Text description of the image to generate |
referenceImages | array | No | Reference images to guide generation. See Reference images |
referenceImages[].fileData | object | No | Reference image supplied as a URL |
referenceImages[].fileData.fileUri | string | Yes | Image URL — scheme must be http, https, or gs |
referenceImages[].fileData.mimeType | string | Yes | MIME type: image/png, image/jpeg, image/webp, image/heic, or image/heif |
referenceImages[].inlineData | object | No | Reference image supplied as base64-encoded data |
referenceImages[].inlineData.data | string | Yes | Base64-encoded image data |
referenceImages[].inlineData.mimeType | string | Yes | MIME type: same five values as fileData.mimeType |
imageConfig | object | No | Image output configuration. Defaults to { "imageSize": "2K" } |
imageConfig.imageSize | string | No | Output resolution: 1K, 2K (default), or 4K, subject to model limits |
imageConfig.aspectRatio | string | No | Aspect ratio. See Aspect ratios for the per-family defaulting rules |
imageConfig.quality | string | No | low, medium, or high. Honoured only by gpt-image-2-official; ignored by every other model |
¹ provider is required on the two generation endpoints and optional on estimate-credits.
² prompt is required on the two generation endpoints; on estimate-credits it may be empty or
omitted (it only affects the input-token estimate).
Each item in referenceImages must carry at least one of fileData or inlineData. Send exactly
one — validation accepts an item with both, but the behaviour is undefined.
provider is validated but not used for routing — the upstream is chosen entirely from model.
Send the vendor that matches your model (as every example on these pages does) so your requests
stay readable, but a mismatch will not change which model runs.
Passing quality to gpt-image-2, gpt-image-2.5-flare, or gpt-image-2.5-sunburst does not
raise an error — the value is dropped and estimate-credits returns the warning
quality_ignored_for_gpt_image_2_lite. Only gpt-image-2-official renders at the requested
quality. See GPT-Image.
Aspect ratios
Fourteen ratios exist across the catalogue. Ten are standard and accepted by most models; four
are extreme and accepted only by gemini-3.1-flash-image and seedream-5-0-pro.
| Ratio | Description | Class |
|---|---|---|
1:1 | Square | Standard |
2:3 | Portrait | Standard |
3:2 | Landscape | Standard |
3:4 | Portrait | Standard |
4:3 | Landscape | Standard |
4:5 | Portrait | Standard |
5:4 | Landscape | Standard |
9:16 | Vertical / mobile | Standard |
16:9 | Widescreen | Standard |
21:9 | Ultra-wide | Standard |
1:4 | Narrow portrait | Extreme |
4:1 | Wide landscape | Extreme |
1:8 | Extreme portrait | Extreme |
8:1 | Panoramic | Extreme |
imageConfig.aspectRatio carries a schema-level default of 1:1. Send an imageConfig object
without an aspectRatio and you get a square image on every model — the per-family behaviour
below only applies when you omit imageConfig entirely, because its own object default is
{ "imageSize": "2K" } and leaves the ratio unset.
How aspectRatio is resolved, once the schema default above has not already filled it in:
| Family | If still unset | If unsupported |
|---|---|---|
| Nano Banana | Resolved to 1:1 | 400, code: 29003 |
| GPT-Image | Left unset — the model picks a ratio for the prompt | 400, code: 26019, with the supported list in the message |
| Seedream | Left unset | Accepted — Seedream is validated on size, not on ratio |
| Wan 2.7 | Left unset | 400, code: 26019, with the supported list in the message |
Image sizes
imageConfig.imageSize accepts 1K, 2K (default), and 4K. Larger sizes cost more credits.
| Model | 1K | 2K | 4K |
|---|---|---|---|
| Every GPT-Image model | Yes | Yes | Yes |
| Both Nano Banana models | Yes | Yes | Yes |
seedream-5-0-pro | Yes | Yes | 400 |
wan2.7-image | Yes | Yes | 400 |
wan2.7-image-pro | Yes | Yes | Text-to-image only — 400 with reference images |
Do not hardcode credit costs — call estimate-credits for the exact figure for your model, size, and ratio.
Reference images
Reference images guide the generation. Supply each one as either a URL (fileData) or base64
(inlineData), never both, and keep the array within the model's limit.
| Family | Max reference images |
|---|---|
| Nano Banana | 14 (the schema cap) |
| GPT-Image | 4 |
| Seedream | 10 |
| Wan 2.7 | 9 |
The schema caps referenceImages at 14 items for every model. Families with a tighter limit reject
the excess with 400 and a message naming the limit. Accepted MIME types for both fileData and
inlineData are image/png, image/jpeg, image/webp, image/heic, and image/heif.
The three image endpoints accept a JSON body up to 25MB — four times the platform default,
because base64 inlineData has to fit inside it. Base64 inflates a file by roughly a third, so keep
inlined images under about 18MB of original bytes in total, or pass URLs with fileData instead.
Generate an image (synchronous)
POST /v1/images/generation
Generate an image and block until it is ready. The response body is the raw model output — JSON
containing base64 image data, not the standard { code, message, data } envelope. See
Response formats.
curl -X POST "https://api.marswave.ai/openapi/v1/images/generation" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"provider": "google",
"model": "gemini-3-pro-image",
"prompt": "An astronaut riding a horse on Mars, photorealistic",
"imageConfig": {
"imageSize": "2K",
"aspectRatio": "16:9"
}
}'const response = await fetch(
'https://api.marswave.ai/openapi/v1/images/generation',
{
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.LISTENHUB_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
provider: 'google',
model: 'gemini-3-pro-image',
prompt: 'An astronaut riding a horse on Mars, photorealistic',
imageConfig: { imageSize: '2K', aspectRatio: '16:9' },
}),
},
)
const result = await response.json()
const image = result.candidates[0].content.parts[0].inlineDataimport os
import requests
response = requests.post(
'https://api.marswave.ai/openapi/v1/images/generation',
headers={'Authorization': f'Bearer {os.environ["LISTENHUB_API_KEY"]}'},
json={
'provider': 'google',
'model': 'gemini-3-pro-image',
'prompt': 'An astronaut riding a horse on Mars, photorealistic',
'imageConfig': {'imageSize': '2K', 'aspectRatio': '16:9'},
},
)
result = response.json()
image = result['candidates'][0]['content']['parts'][0]['inlineData']With reference images
Reference images work the same way on every endpoint. Use fileData when the image is already
hosted, and inlineData when you have the bytes in memory.
curl -X POST "https://api.marswave.ai/openapi/v1/images/generation" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"provider": "google",
"model": "gemini-3-pro-image",
"prompt": "Redraw this character in a watercolor style",
"referenceImages": [
{
"fileData": {
"fileUri": "https://example.com/character.png",
"mimeType": "image/png"
}
}
],
"imageConfig": { "imageSize": "2K", "aspectRatio": "3:4" }
}'curl -X POST "https://api.marswave.ai/openapi/v1/images/generation" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"provider": "google",
"model": "gemini-3-pro-image",
"prompt": "Redraw this character in a watercolor style",
"referenceImages": [
{
"inlineData": {
"data": "'"$(base64 -w0 character.png)"'",
"mimeType": "image/png"
}
}
],
"imageConfig": { "imageSize": "2K", "aspectRatio": "3:4" }
}'Requests carrying base64 reference images hit a separate, tighter rate limit. On a 429, read the
Retry-After header and back off before retrying.
Asynchronous generation
For long-running or high-resolution requests, submit a task and poll for the result instead of holding a request open.
Create an async task
POST /v1/images/generation/async
Same request body as the synchronous endpoint. Returns 202 with a taskId immediately; generation
runs in the background and the resulting images are persisted as hosted URLs.
curl -X POST "https://api.marswave.ai/openapi/v1/images/generation/async" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"provider": "google",
"model": "gemini-3-pro-image",
"prompt": "An astronaut riding a horse on Mars, photorealistic",
"imageConfig": { "imageSize": "4K", "aspectRatio": "16:9" }
}'Response (202):
{
"code": 0,
"message": "",
"data": { "taskId": "65f0...", "status": "pending" }
}Get a task
GET /v1/images/generation/tasks/{taskId}
Poll for the status and result of a single task. status is one of pending, generating,
success, or fail. On success, images holds the hosted result URLs.
curl "https://api.marswave.ai/openapi/v1/images/generation/tasks/65f0abc..." \
-H "Authorization: Bearer $LISTENHUB_API_KEY"{
"code": 0,
"message": "",
"data": {
"taskId": "65f0abc...",
"status": "success",
"images": [{ "url": "https://.../0.png", "mimeType": "image/png" }],
"createdAt": 1750000000000,
"completedAt": 1750000020000
}
}| Field | Type | Description |
|---|---|---|
taskId | string | Task identifier |
status | string | pending, generating, success, or fail |
images | array | Present on success; each item is { url, mimeType } |
failMsg | string | Failure message when status is fail |
createdAt | number | Creation time (epoch milliseconds) |
completedAt | number | Completion time (epoch milliseconds), when finished |
List tasks
GET /v1/images/generation/tasks
List your image tasks, newest first.
| Query param | Type | Default | Description |
|---|---|---|---|
page | number | 1 | Page number, minimum 1 |
pageSize | number | 20 | Items per page, 1–100 |
status | string | — | Filter by pending, generating, success, or fail |
curl "https://api.marswave.ai/openapi/v1/images/generation/tasks?page=1&pageSize=20&status=success" \
-H "Authorization: Bearer $LISTENHUB_API_KEY"{
"code": 0,
"message": "",
"data": {
"items": [
{
"taskId": "65f0...",
"status": "success",
"images": [],
"createdAt": 1750000000000,
"completedAt": 1750000020000
}
],
"page": 1,
"pageSize": 20,
"total": 1
}
}Background tasks left in pending or generating for more than 30 minutes are swept to fail
with a timeout failMsg. Treat a non-terminal status older than that as failed and retry.
Estimate credits
POST /v1/images/generation/estimate-credits
Returns the credit cost for a given configuration and whether your account can afford it, without
spending credits or calling the model. Accepts the same body as the generation endpoints;
provider and prompt are optional here.
curl -X POST "https://api.marswave.ai/openapi/v1/images/generation/estimate-credits" \
-H "Authorization: Bearer $LISTENHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"imageConfig": {
"imageSize": "2K",
"aspectRatio": "1:1"
}
}'const response = await fetch(
'https://api.marswave.ai/openapi/v1/images/generation/estimate-credits',
{
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.LISTENHUB_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'gpt-image-2',
imageConfig: { imageSize: '2K', aspectRatio: '1:1' },
}),
},
)
const { data } = await response.json()
console.log(data.credits, data.canGenerate)import os
import requests
response = requests.post(
'https://api.marswave.ai/openapi/v1/images/generation/estimate-credits',
headers={'Authorization': f'Bearer {os.environ["LISTENHUB_API_KEY"]}'},
json={
'model': 'gpt-image-2',
'imageConfig': {'imageSize': '2K', 'aspectRatio': '1:1'},
},
)
data = response.json()['data']
print(data['credits'], data['canGenerate'])Estimate response
Wrapped in the standard envelope. The data object:
| Field | Type | Description |
|---|---|---|
model | string | Normalized GA model ID used for the estimate |
imageSize | string | Resolved output size |
aspectRatio | string | Resolved aspect ratio (absent when the model auto-selects) |
quality | string | Resolved quality (absent when not applicable) |
pixels | object | { "width": number, "height": number, "size": "WxH" } when resolved |
credits | number | Credits this configuration would cost |
canGenerate | boolean | Whether the account has enough effective credit balance |
requiresSubscription | boolean | true only for gpt-image-2-official at 4K or quality: high |
pricing | object | Pricing metadata: pricingVersion, mode (token-estimate or fixed), and rate inputs |
warnings | array | Advisory strings, e.g. quality_ignored_for_gpt_image_2_lite |
{
"code": 0,
"message": "",
"data": {
"model": "gpt-image-2",
"imageSize": "2K",
"aspectRatio": "1:1",
"credits": 6,
"canGenerate": true,
"requiresSubscription": false,
"pricing": { "pricingVersion": "gpt-image-2-lite-fixed-2026-06-15", "mode": "fixed" },
"warnings": []
}
}Most models are priced at a fixed credit cost per size (mode: "fixed"). gpt-image-2-official
is priced from estimated tokens (mode: "token-estimate"), so its cost moves with prompt length,
output dimensions, requested quality, and the number of reference images (each counted at a flat
rate over the API).
credits is the credit price of the configuration, always. It does not subtract free quota —
a 1K or 2K request on an account that still has quota estimates at full price and then costs
nothing. Read freeUsages yourself if you need to
show the price the account will actually pay.
Free generations
API key calls draw on the same account-level free-quota (freeUsages) balance as the ListenHub and
Labnana apps. You earn quota through sign-up, invites, and check-ins, then spend it through the API.
Query the live balance via
GET /v1/user/subscription and read its freeUsages
map.
When the matching balance is greater than 0, a 1K or 2K request spends one free generation
instead of credits. Once the balance hits 0, the same request falls back to normal credit billing.
| Free-quota resource | Models it covers |
|---|---|
gemini-3-pro-image-relax-1k-2k | gemini-3-pro-image |
gemini-3.1-flash-image | gemini-3.1-flash-image |
gpt-image-2 | gpt-image-2 |
gpt-image-2.5 | gpt-image-2.5-flare, gpt-image-2.5-sunburst |
seedream-5-0-pro | seedream-5-0-pro |
wan2.7-image | wan2.7-image |
Free quota applies only to 1K and 2K sizes — a 4K request never draws from freeUsages
and is always billed in credits. gpt-image-2-official and wan2.7-image-pro have no free quota
at any size.
gemini-3-pro-image is the only model whose free quota runs on a separate execution lane, with its
own failure metadata to branch on — see
Nano Banana → the relax lane.
Response formats
| Endpoint | Wrapped envelope? | Body |
|---|---|---|
POST /v1/images/generation | No — raw model JSON | Base64 image data (see below) |
POST /v1/images/generation/async | Yes | { taskId, status } |
POST /v1/images/generation/estimate-credits | Yes | Estimate object |
GET /v1/images/generation/tasks | Yes | Paginated { items, page, pageSize, total } |
GET /v1/images/generation/tasks/{taskId} | Yes | Task object |
The synchronous endpoint returns the raw model output directly. A successful body contains the generated image as base64 data:
{
"candidates": [
{
"content": {
"parts": [
{
"inlineData": {
"mimeType": "image/png",
"data": "<BASE64_ENCODED_IMAGE>"
}
}
]
}
}
]
}Decode the data field from base64 to obtain the image file. The async path persists the image
for you and returns hosted urls on the task object — no base64 decoding needed.
Error codes
Errors use the standard envelope with a non-zero code or, for the synchronous endpoint, the raw
error body described in Free generations.
These routes normalize every business error to HTTP 400. Insufficient credits is a 400 with
code: 26004, not a 402 — branch on the envelope's code, not on the HTTP status. 429 is the
one status that carries its own meaning. The exception is the synchronous free-relax failure body,
which is raw rather than enveloped and carries errno instead of code.
| HTTP status | code | Meaning |
|---|---|---|
400 | 29003 | Invalid parameters — a size or reference-image count the model rejects, or a ratio the Nano Banana family rejects |
400 | 26019 | Unsupported aspect ratio on GPT-Image or Wan 2.7 |
400 | 26004 | Insufficient credits |
400 | 23016 | Subscription required — gpt-image-2-official at 4K or quality: high on an account with no active plan. Check requiresSubscription on the estimate first |
429 | — | Rate limited or service busy — read Retry-After and back off |
Ratio rejections do not all carry one code: Nano Banana raises 29003, while GPT-Image and Wan 2.7
raise 26019. Branch on both if you are detecting "wrong ratio for this model".
Both generation endpoints are rate limited per user and globally. Requests carrying base64 reference
images (inlineData) are subject to a separate, tighter limit and may be throttled first at peak
times. estimate-credits is not rate limited.