POST /v1/images/generations
OpenAI Images API: base64 in the body, all-or-nothing.
Generate images from a text prompt. This is the OpenAI Images API surface: point client.images.generate at https://api.routerplus.com/v1 and the call works unchanged. The path https://api.routerplus.com/images/generations is an alias for the same handler, for SDKs that omit the /v1 prefix.
To try it without writing code, pick an image model in the playground.
The route is deliberately small: generations only (no edits, no variations), non-streaming, and the image always comes back as base64 in the body. Everything else is a typed error, never a silent degradation. The GPT Image models are served by OpenAI direct. Google's Nano Banana family, ByteDance's Seedream and xAI's Grok Imagine are served through OpenRouter today and billed at the cost it reports — see Models billed at reported cost.
Authentication
| Header | Format |
|---|---|
Authorization | Bearer tm_vk_... (the standard OpenAI-style header) |
x-api-key | tm_vk_... — also accepted, same key |
A missing or invalid key returns 401 with error_type: "auth". Optionally send HTTP-Referer and X-Title to identify your app (content-free attribution).
Request
JSON body, capped at 10 MB (413 request_too_large above that). Every rule below is checked before anything is reserved or dispatched, so a refused request has no ledger row, no x-tm-attempts header, and no charge.
| Parameter | Type | Behavior |
|---|---|---|
model | string, required | An image model id — see Model discovery. A dated pin (-20260421, -2026-04-21, -latest) resolves to its catalog family for routing and pricing, and your exact string goes upstream. A chat model here is a 400 naming the chat routes; a video model a 400 naming /v1/videos. An id that is not an image model anywhere is a 404. Omitted or empty is a 400 naming model. |
prompt | string, required | Non-empty, at most 32 000 characters. |
n | integer | 1 to 4 on the GPT Image models; the OpenRouter models take 1 (Seedream 5.0 Lite up to 4). Defaults to 1. Ours is stricter than the SDK's 10: four buffered images already approach the response cap. |
size | string | auto, 1024x1024, 1536x1024 or 1024x1536 on the GPT Image models. The OpenRouter models take aspect_ratio instead, and drop size. |
aspect_ratio | string | On the OpenRouter models: 1:1, 16:9, 9:16, … — each model's own list. |
resolution | string | On most OpenRouter models: 512, 1K, 2K or 4K, per model. |
seed | integer | On Seedream. |
quality | string | auto, low, medium, high — plus xhigh and max on gpt-image-2.5-flare and gpt-image-2.5-sunburst; low or medium on Grok Imagine. |
background | string | auto, transparent, opaque on the GPT Image models. |
moderation | string | auto or low on the GPT Image models. |
output_format | string | png, jpeg or webp on the GPT Image models. |
output_compression | integer | 0 to 100 on the GPT Image models. |
user | string | Forwarded verbatim. Anything but a string is a 400. |
response_format | string | b64_json only, and it is a no-op. url is a 400 — see Always base64. |
stream, partial_images | — | Absent, false and 0 are no-ops. Anything else is a 400: this route does not stream. |
provider | object | Routing controls read by the gateway and never forwarded — require_parameters and the rest, see Routing policies. |
The accepted values for size, aspect_ratio, resolution, quality, background, moderation, output_format, output_compression and seed are per model and published in /api/models.json under supported_parameters. A value outside a model's declared set is a 400 whose message starts with the field name and lists what that model accepts. One of these parameters that a model does not declare at all — size on an OpenRouter model, quality on Nano Banana — is dropped and recorded, like any unknown key.
An explicit null means "not provided" for every optional parameter above, exactly like omitting the key: the OpenAI SDKs serialise an unset optional as null, so client.images.generate(model=…, prompt=…, n=None, size=None) works. A null is never forwarded upstream and never recorded as a drop. model and prompt are required, so null there is still a 400.
Everything else
style and any key not in the table are dropped and recorded, never forwarded and never guessed at (D8 §2). The dropped names are comma-joined in the x-tm-dropped-params response header, stored on the attempt's ledger row, and visible at GET /v1/generation?id=. To refuse instead of dropping, send "provider": {"require_parameters": true} — a would-be drop then becomes a 400 before any money is reserved. provider itself is consumed by the gateway and never forwarded.
Always base64
The gateway always returns the image bytes inline as b64_json. There is no hosted URL: storing an image would make this service stateful and would put buyer content on our disks, which the data policy forbids. response_format: "url" is therefore a typed 400, not a quietly different answer. Decode the base64 yourself — the examples below do it in three languages.
Response
HTTP 200, application/json, with OpenAI's own ImagesResponse relayed untouched except for one added field:
{
"created": 1790000000,
"data": [{ "b64_json": "iVBORw0KGgoAAAANSUhEUg..." }],
"background": "opaque",
"output_format": "png",
"quality": "medium",
"size": "1024x1024",
"usage": {
"input_tokens": 12,
"input_tokens_details": { "text_tokens": 12, "image_tokens": 0 },
"output_tokens": 1056,
"total_tokens": 1068,
"cost": 0.0423
}
}| Field | Meaning |
|---|---|
data[].b64_json | The image, base64-encoded. One entry per image the provider rendered — normally n; the gateway relays what it received and bills what was reported. |
data[].revised_prompt | The provider's rewritten prompt, when it sends one. Relayed, never stored. |
usage.input_tokens | Prompt tokens, billed at the model's prompt rate. |
usage.output_tokens | Image output tokens, billed at the model's completion rate. |
usage.input_tokens_details | Informational only — the bill is computed from the two counts above. |
usage.total_tokens | input_tokens + output_tokens. |
usage.cost | USD, the full request debit including any attempt that failed over before this one. 0 on BYOK. Absent on static dev keys. |
The chat surface's prompt_tokens_details is never injected here; an images response keeps its own shape.
If the provider returns a 200 without a usage object, or with input_tokens and output_tokens both 0, the gateway bills n × the model's per-image ceiling with provenance estimated — the same amount the reservation held — and writes those numbers into usage so your copy of the bill still adds up. The provenance is visible at GET /v1/generation.
Billing
Same two SKUs as chat, same integer micro-USD math (Pricing & billing): text in bills at prompt, image out bills at completion.
- Reserve, before any provider is contacted: prompt bytes ÷ 4 at the
promptrate, plusn× the model's per-image ceiling at thecompletionrate, rounded up. The chat path'smax_tokensdefault plays no part here. - Settle, on the provider-reported counts, rounded down. The hold is released in full.
Worked numbers on gpt-image-1 (ceiling 6240 tokens, $5 / $40 per million):
| Case | Amount |
|---|---|
Reserve, n: 1 | 6240 × $40/M ≈ $0.2496, plus the prompt estimate |
Reserve, n: 4 | 4 × 6240 × $40/M ≈ $0.9984, plus the prompt estimate |
| Settle, one medium 1024×1024 image | 1056 × $40/M = $0.04224, plus the prompt |
All-or-nothing. A generation either arrives whole or costs nothing. A failed attempt, a response over the 32 MiB cap, and a 200 that carries no decodable image all settle at zero. Once a 200 has been received the gateway never re-dispatches, so one request can never buy two renders.
A cancelled render is billed at what the provider reported. If you hang up mid-render the upstream call is not aborted: the render finishes, the usage is read, and the attempt settles the provider's own numbers — never the reservation, never a pretend $0. Nothing is written to your dead socket.
Reservations are large next to a trial balance. n: 4 at a high quality holds about a dollar, and your balance must cover every reservation your org has in flight. A 429 insufficient_quota on n: 4 is correct behavior, not a bug — lower n, or add credits.
Errors
| Case | Status | error_type | Billed |
|---|---|---|---|
A field outside the model's accepted values, n outside 1–4, stream, partial_images, response_format: "url", a chat or video model on this route | 400 | invalid_request | no row |
| The model is listed but has no price row | 400 | model_not_priced | no row |
| Not an image model in the catalog | 404 | model_unavailable | no row |
| Request body over 10 MB | 413 | request_too_large | no row |
| Key RPM or a shared limit, balance too low, or a monthly spend cap | 429 | rate_limit / insufficient_quota | no row |
OpenAI moderation refusal (moderation_blocked) | the provider's, usually 400 | content_policy | $0; never rerouted, never a health strike |
| Any other provider 4xx | the provider's | upstream_error | $0; no failover |
| Provider 5xx, a failed connection, or the 450 s budget | 5xx relayed / 502 / 504 | upstream_error / model_unavailable | $0 per failed attempt; fails over like chat |
| Response over the 32 MiB cap | 502 | gateway_error | $0; no failover |
A 200 that is unparseable or has no b64_json | 502 | upstream_error | $0; no failover |
| The budget ran out while reading the image body | 504 | upstream_error | $0; no failover |
| Every deployment serving the model is cooling down | 503 | gateway_error | nothing dispatched; retry-after: 5 |
| The ledger journal, the shared limit store, or the admission lease is unavailable | 503 | gateway_error | nothing charged |
Bodies follow the OpenAI error envelope with the canonical class in error.metadata.error_type, and the class is on every failure in x-tm-error-code. A provider's own error body is never relayed; its status is in x-tm-upstream-status. The full taxonomy and remediation table is in Errors.
The reverse refusal also holds: an image model sent to POST /v1/chat/completions, POST /v1/messages or POST /v1/videos is a 400 invalid_request naming /v1/images/generations.
Limits
| Limit | Value |
|---|---|
| Images per request | n ≤ 4 |
| Prompt length | 32 000 characters |
| Request body | 10 MB → 413 request_too_large |
| Upstream response | 32 MiB → 502 gateway_error, nothing billed. A 4K Nano Banana 2 image is about 20 MB of base64. |
| Generation budget | 450 s, covering the body read as well as the render |
| Streaming | not supported |
| Admission | one slot held for the whole render (typically 10–60 s) |
| Providers | OpenAI direct for the GPT Image models, OpenRouter for the rest — never Azure, Bedrock or an Anthropic deployment |
Models billed at reported cost
The OpenRouter models are billed at the cost OpenRouter reports for the render, passed through with no margin (Pricing & billing).
- Reserve:
n× the model's per-image ceiling, in USD — Per image ≤ on the models page. - Settle: the reported cost, exactly. None reported: the ceiling, marked
estimated. - The response keeps OpenAI's shape.
usage.input_tokensandusage.output_tokensare OpenRouter's own counts, for reading only;usage.costis the charge.
| Model id | Listed price | Per image ≤ | Measured (2026-09-23) |
|---|---|---|---|
gemini-2.5-flash-image (Nano Banana) | $30 / M image tokens | $0.05 | $0.0387, 5–7 s |
gemini-3.1-flash-image (Nano Banana 2) | $60 / M image tokens | $0.20 | $0.0448 at 512 to $0.1512 at 4K, 7–30 s |
gemini-3-pro-image (Nano Banana Pro) | $120 / M image tokens | $0.30 | $0.1344 at 1K and 2K, $0.2409 at 4K, 18–34 s |
seedream-5.0-pro | $0.045; $0.09 at 2K | $0.10 | as listed, 40–67 s |
seedream-5.0-lite | $0.035 | $0.04 | as listed, 14–30 s |
grok-imagine-image-2.0 | $0.04–$0.08 by quality and resolution | $0.09 | $0.06 at 1K, $0.08 at 2K (medium), 32–46 s |
Response headers
| Header | When | Meaning |
|---|---|---|
x-request-id | every call | the id GET /v1/generation?id= audits |
x-tm-provider | served requests | deployment that rendered the image |
x-tm-attempts | after ≥ 1 dispatch | physical dispatches, failovers included |
x-tm-upstream-status | when a provider answered | the provider's own HTTP status |
x-tm-dropped-params | when non-empty | comma-joined names of parameters the gateway stripped |
x-tm-upstream-model | aggregator swaps only | the id the gateway actually sent |
x-tm-served-by | aggregator routes that name it | the provider the aggregator used |
x-tm-route-plan-id | billed requests | the route plan that chose the deployment |
x-tm-admission-mode, x-tm-remaining-rpm, x-tm-remaining-tpm | billed requests | how the request was admitted and the headroom left — see Limits and capacity |
x-tm-error-code | every failure | the canonical error class |
x-tm-error-origin, x-tm-limit-scope, x-tm-limit-kind | limit, provider and infrastructure failures | where it came from, and the limit that refused it |
There is no image-specific header. Browsers can read all of these except x-tm-error-code and x-tm-route-plan-id, which are not CORS-exposed.
Examples
curl -s https://api.routerplus.com/v1/images/generations \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"gpt-image-1","prompt":"a red bicycle on a white background","size":"1024x1024","quality":"medium"}' \
| jq -r '.data[0].b64_json' | base64 -d > out.pngopenai SDK with the base URL swappedimport base64
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
r = client.images.generate(
model="gpt-image-1",
prompt="a red bicycle on a white background",
size="1024x1024",
quality="medium",
)
with open("out.png", "wb") as f:
f.write(base64.b64decode(r.data[0].b64_json))
print(r.usage.model_dump()["cost"]) # the exact ledger debit, in USDopenai packageimport { writeFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY });
const r = await client.images.generate({
model: "gpt-image-1",
prompt: "a red bicycle on a white background",
size: "1024x1024",
quality: "medium",
});
const b64 = r.data?.[0]?.b64_json ?? "";
writeFileSync("out.png", Buffer.from(b64, "base64"));
console.log(r.usage); // input/output tokens and cost (USD)Content-free
Prompts and image bytes are content, and the marketplace never stores content. The prompt is never logged and never written to the ledger; the base64 bytes and any revised_prompt are relayed to you and dropped. Telemetry keeps the response byte size so throughput stays measurable, and nothing else. See Data policy.
Model discovery
curl -s "https://api.routerplus.com/v1/models?output_modalities=image" \
-H "Authorization: Bearer $TM_API_KEY"Every row of GET /v1/models carries architecture.output_modalities — ["text"], ["image"] or ["video"] — and the ?output_modalities= filter takes a comma list of those values. The public feed https://app.routerplus.com/api/models.json carries output_modalities, the full supported_parameters descriptor map for image models, and max_output_tokens as the per-image ceiling. See GET /v1/models and Models & catalog.
BYOK
Image generation works through an openai provider connection. List the exact image ids you intend to send on the connection — matching is literal, and an id the connection does not carry is a 404 model_unavailable. The model's metadata (ceiling, descriptors) still comes from the house catalog. usage.cost is 0, the marketplace charge is zero, and OpenAI bills your own account for the render. See Bring your own key.
See also: Models & catalog · Pricing & billing · Errors · Wire compatibility
Markdown source for agents: /docs/api-images.md · index at /llms.txt