# POST /v1/images/generations

Generate images from a text prompt. This is the OpenAI Images API surface: point
`client.images.generate` at `https://api.routerplus.com/v1` and the call works unchanged. The path
`https://api.routerplus.com/images/generations` is an alias for the same handler, for SDKs that omit
the `/v1` prefix.

To try it without writing code, pick an image model in the [playground](/docs/playground).

The route is deliberately small: generations only (no edits, no variations), non-streaming,
and the image always comes back as base64 in the body. Everything else is a typed error,
never a silent degradation. The GPT Image models are served by OpenAI direct. Google's
Nano Banana family, ByteDance's Seedream and xAI's Grok Imagine are served through
OpenRouter today and billed at the cost it reports — see
[Models billed at reported cost](#models-billed-at-reported-cost).

## Authentication

| Header | Format |
|---|---|
| `Authorization` | `Bearer tm_vk_...` (the standard OpenAI-style header) |
| `x-api-key` | `tm_vk_...` — also accepted, same key |

A missing or invalid key returns 401 with `error_type: "auth"`. Optionally send
`HTTP-Referer` and `X-Title` to identify your app (content-free attribution).

## Request

JSON body, capped at 10 MB (413 `request_too_large` above that). Every rule below is
checked **before** anything is reserved or dispatched, so a refused request has no ledger
row, no `x-tm-attempts` header, and no charge.

| Parameter | Type | Behavior |
|---|---|---|
| `model` | string, required | An image model id — see [Model discovery](#model-discovery). A dated pin (`-20260421`, `-2026-04-21`, `-latest`) resolves to its catalog family for routing and pricing, and your exact string goes upstream. A **chat** model here is a 400 naming the chat routes; a video model a 400 naming `/v1/videos`. An id that is not an image model anywhere is a 404. Omitted or empty is a 400 naming `model`. |
| `prompt` | string, required | Non-empty, at most 32 000 characters. |
| `n` | integer | 1 to 4 on the GPT Image models; the OpenRouter models take 1 (Seedream 5.0 Lite up to 4). Defaults to 1. Ours is stricter than the SDK's 10: four buffered images already approach the response cap. |
| `size` | string | `auto`, `1024x1024`, `1536x1024` or `1024x1536` on the GPT Image models. The OpenRouter models take `aspect_ratio` instead, and drop `size`. |
| `aspect_ratio` | string | On the OpenRouter models: `1:1`, `16:9`, `9:16`, … — each model's own list. |
| `resolution` | string | On most OpenRouter models: `512`, `1K`, `2K` or `4K`, per model. |
| `seed` | integer | On Seedream. |
| `quality` | string | `auto`, `low`, `medium`, `high` — plus `xhigh` and `max` on `gpt-image-2.5-flare` and `gpt-image-2.5-sunburst`; `low` or `medium` on Grok Imagine. |
| `background` | string | `auto`, `transparent`, `opaque` on the GPT Image models. |
| `moderation` | string | `auto` or `low` on the GPT Image models. |
| `output_format` | string | `png`, `jpeg` or `webp` on the GPT Image models. |
| `output_compression` | integer | 0 to 100 on the GPT Image models. |
| `user` | string | Forwarded verbatim. Anything but a string is a 400. |
| `response_format` | string | `b64_json` only, and it is a no-op. `url` is a 400 — see [Always base64](#always-base64). |
| `stream`, `partial_images` | — | Absent, `false` and `0` are no-ops. Anything else is a 400: this route does not stream. |
| `provider` | object | Routing controls read by the gateway and never forwarded — `require_parameters` and the rest, see [Routing policies](/docs/routing-policies). |

The accepted values for `size`, `aspect_ratio`, `resolution`, `quality`, `background`,
`moderation`, `output_format`, `output_compression` and `seed` are **per model** and
published in `/api/models.json` under `supported_parameters`. A value outside a model's
declared set is a 400 whose message starts with the field name and lists what that model
accepts. One of these parameters that a model does not declare at all — `size` on an
OpenRouter model, `quality` on Nano Banana — is dropped and recorded, like any unknown key.

An explicit `null` means "not provided" for every optional parameter above, exactly like
omitting the key: the OpenAI SDKs serialise an unset optional as `null`, so
`client.images.generate(model=…, prompt=…, n=None, size=None)` works. A `null` is never
forwarded upstream and never recorded as a drop. `model` and `prompt` are required, so
`null` there is still a 400.

### Everything else

`style` and any key not in the table are **dropped and recorded**, never forwarded and
never guessed at (D8 §2). The dropped names are comma-joined in the `x-tm-dropped-params`
response header, stored on the attempt's ledger row, and visible at
`GET /v1/generation?id=`. To refuse instead of dropping, send
`"provider": {"require_parameters": true}` — a would-be drop then becomes a 400 before any
money is reserved. `provider` itself is consumed by the gateway and never forwarded.

### Always base64

The gateway always returns the image bytes inline as `b64_json`. There is no hosted URL:
storing an image would make this service stateful and would put buyer content on our
disks, which the [data policy](/docs/data-policy) forbids. `response_format: "url"` is
therefore a typed 400, not a quietly different answer. Decode the base64 yourself — the
examples below do it in three languages.

## Response

HTTP 200, `application/json`, with OpenAI's own `ImagesResponse` relayed untouched except
for one added field:

```json
{
  "created": 1790000000,
  "data": [{ "b64_json": "iVBORw0KGgoAAAANSUhEUg..." }],
  "background": "opaque",
  "output_format": "png",
  "quality": "medium",
  "size": "1024x1024",
  "usage": {
    "input_tokens": 12,
    "input_tokens_details": { "text_tokens": 12, "image_tokens": 0 },
    "output_tokens": 1056,
    "total_tokens": 1068,
    "cost": 0.0423
  }
}
```

| Field | Meaning |
|---|---|
| `data[].b64_json` | The image, base64-encoded. One entry per image the provider rendered — normally `n`; the gateway relays what it received and bills what was reported. |
| `data[].revised_prompt` | The provider's rewritten prompt, when it sends one. Relayed, never stored. |
| `usage.input_tokens` | Prompt tokens, billed at the model's `prompt` rate. |
| `usage.output_tokens` | Image output tokens, billed at the model's `completion` rate. |
| `usage.input_tokens_details` | Informational only — the bill is computed from the two counts above. |
| `usage.total_tokens` | `input_tokens + output_tokens`. |
| `usage.cost` | USD, the **full request debit** including any attempt that failed over before this one. `0` on BYOK. Absent on static dev keys. |

The chat surface's `prompt_tokens_details` is never injected here; an images response keeps
its own shape.

> [!NOTE]
> If the provider returns a 200 without a usage object, or with `input_tokens` and
> `output_tokens` both 0, the gateway bills `n` × the model's
> per-image ceiling with provenance `estimated` — the same amount the reservation held —
> and writes those numbers into `usage` so your copy of the bill still adds up. The
> provenance is visible at [GET /v1/generation](/docs/api-usage).

## Billing

Same two SKUs as chat, same integer micro-USD math ([Pricing & billing](/docs/pricing)):
text in bills at `prompt`, image out bills at `completion`.

- **Reserve**, before any provider is contacted: prompt bytes ÷ 4 at the `prompt` rate,
  plus `n` × the model's per-image ceiling at the `completion` rate, rounded up. The chat
  path's `max_tokens` default plays no part here.
- **Settle**, on the provider-reported counts, rounded down. The hold is released in full.

Worked numbers on `gpt-image-1` (ceiling 6240 tokens, $5 / $40 per million):

| Case | Amount |
|---|---|
| Reserve, `n: 1` | 6240 × $40/M ≈ $0.2496, plus the prompt estimate |
| Reserve, `n: 4` | 4 × 6240 × $40/M ≈ $0.9984, plus the prompt estimate |
| Settle, one medium 1024×1024 image | 1056 × $40/M = $0.04224, plus the prompt |

**All-or-nothing.** A generation either arrives whole or costs nothing. A failed attempt,
a response over the 32 MiB cap, and a 200 that carries no decodable image all settle at
zero. Once a 200 has been received the gateway never re-dispatches, so one request can
never buy two renders.

**A cancelled render is billed at what the provider reported.** If you hang up mid-render
the upstream call is *not* aborted: the render finishes, the usage is read, and the attempt
settles the provider's own numbers — never the reservation, never a pretend $0. Nothing is
written to your dead socket.

> [!WARNING]
> Reservations are large next to a trial balance. `n: 4` at a high quality holds about a
> dollar, and your balance must cover every reservation your org has in flight. A
> `429 insufficient_quota` on `n: 4` is correct behavior, not a bug — lower `n`, or add
> credits.

## Errors

| Case | Status | `error_type` | Billed |
|---|---|---|---|
| A field outside the model's accepted values, `n` outside 1–4, `stream`, `partial_images`, `response_format: "url"`, a chat or video model on this route | 400 | `invalid_request` | no row |
| The model is listed but has no price row | 400 | `model_not_priced` | no row |
| Not an image model in the catalog | 404 | `model_unavailable` | no row |
| Request body over 10 MB | 413 | `request_too_large` | no row |
| Key RPM or a shared limit, balance too low, or a monthly spend cap | 429 | `rate_limit` / `insufficient_quota` | no row |
| OpenAI moderation refusal (`moderation_blocked`) | the provider's, usually 400 | `content_policy` | $0; never rerouted, never a health strike |
| Any other provider 4xx | the provider's | `upstream_error` | $0; no failover |
| Provider 5xx, a failed connection, or the 450 s budget | 5xx relayed / 502 / 504 | `upstream_error` / `model_unavailable` | $0 per failed attempt; fails over like chat |
| Response over the 32 MiB cap | 502 | `gateway_error` | $0; no failover |
| A 200 that is unparseable or has no `b64_json` | 502 | `upstream_error` | $0; no failover |
| The budget ran out while reading the image body | 504 | `upstream_error` | $0; no failover |
| Every deployment serving the model is cooling down | 503 | `gateway_error` | nothing dispatched; `retry-after: 5` |
| The ledger journal, the shared limit store, or the admission lease is unavailable | 503 | `gateway_error` | nothing charged |

Bodies follow the OpenAI error envelope with the canonical class in
`error.metadata.error_type`, and the class is on every failure in `x-tm-error-code`. A
provider's own error body is never relayed; its status is in `x-tm-upstream-status`. The
full taxonomy and remediation table is in [Errors](/docs/errors).

The reverse refusal also holds: an image model sent to `POST /v1/chat/completions`,
`POST /v1/messages` or `POST /v1/videos` is a 400 `invalid_request` naming
`/v1/images/generations`.

## Limits

| Limit | Value |
|---|---|
| Images per request | `n` ≤ 4 |
| Prompt length | 32 000 characters |
| Request body | 10 MB → 413 `request_too_large` |
| Upstream response | 32 MiB → 502 `gateway_error`, nothing billed. A 4K Nano Banana 2 image is about 20 MB of base64. |
| Generation budget | 450 s, covering the body read as well as the render |
| Streaming | not supported |
| Admission | one slot held for the whole render (typically 10–60 s) |
| Providers | OpenAI direct for the GPT Image models, OpenRouter for the rest — never Azure, Bedrock or an Anthropic deployment |

## Models billed at reported cost

The OpenRouter models are billed at **the cost OpenRouter reports** for the render,
passed through with no margin ([Pricing & billing](/docs/pricing#models-billed-at-the-providers-reported-cost)).

- **Reserve:** `n` × the model's per-image ceiling, in USD — **Per image ≤** on the models
  page.
- **Settle:** the reported cost, exactly. None reported: the ceiling, marked `estimated`.
- **The response** keeps OpenAI's shape. `usage.input_tokens` and `usage.output_tokens`
  are OpenRouter's own counts, for reading only; `usage.cost` is the charge.

| Model id | Listed price | Per image ≤ | Measured (2026-09-23) |
|---|---|---|---|
| `gemini-2.5-flash-image` (Nano Banana) | $30 / M image tokens | $0.05 | $0.0387, 5–7 s |
| `gemini-3.1-flash-image` (Nano Banana 2) | $60 / M image tokens | $0.20 | $0.0448 at 512 to $0.1512 at 4K, 7–30 s |
| `gemini-3-pro-image` (Nano Banana Pro) | $120 / M image tokens | $0.30 | $0.1344 at 1K and 2K, $0.2409 at 4K, 18–34 s |
| `seedream-5.0-pro` | $0.045; $0.09 at 2K | $0.10 | as listed, 40–67 s |
| `seedream-5.0-lite` | $0.035 | $0.04 | as listed, 14–30 s |
| `grok-imagine-image-2.0` | $0.04–$0.08 by quality and resolution | $0.09 | $0.06 at 1K, $0.08 at 2K (medium), 32–46 s |

## Response headers

| Header | When | Meaning |
|---|---|---|
| `x-request-id` | every call | the id `GET /v1/generation?id=` audits |
| `x-tm-provider` | served requests | deployment that rendered the image |
| `x-tm-attempts` | after ≥ 1 dispatch | physical dispatches, failovers included |
| `x-tm-upstream-status` | when a provider answered | the provider's own HTTP status |
| `x-tm-dropped-params` | when non-empty | comma-joined names of parameters the gateway stripped |
| `x-tm-upstream-model` | aggregator swaps only | the id the gateway actually sent |
| `x-tm-served-by` | aggregator routes that name it | the provider the aggregator used |
| `x-tm-route-plan-id` | billed requests | the route plan that chose the deployment |
| `x-tm-admission-mode`, `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | billed requests | how the request was admitted and the headroom left — see [Limits and capacity](/docs/admission) |
| `x-tm-error-code` | every failure | the canonical error class |
| `x-tm-error-origin`, `x-tm-limit-scope`, `x-tm-limit-kind` | limit, provider and infrastructure failures | where it came from, and the limit that refused it |

There is no image-specific header. Browsers can read all of these except `x-tm-error-code`
and `x-tm-route-plan-id`, which are not CORS-exposed.

## Examples

curl — decode the first image straight to a file:

```bash
curl -s https://api.routerplus.com/v1/images/generations \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-image-1","prompt":"a red bicycle on a white background","size":"1024x1024","quality":"medium"}' \
  | jq -r '.data[0].b64_json' | base64 -d > out.png
```

Python — the official `openai` SDK with the base URL swapped:

```python
import base64
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

r = client.images.generate(
    model="gpt-image-1",
    prompt="a red bicycle on a white background",
    size="1024x1024",
    quality="medium",
)
with open("out.png", "wb") as f:
    f.write(base64.b64decode(r.data[0].b64_json))
print(r.usage.model_dump()["cost"])  # the exact ledger debit, in USD
```

TypeScript — the official `openai` package:

```typescript
import { writeFileSync } from "node:fs";
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY });

const r = await client.images.generate({
  model: "gpt-image-1",
  prompt: "a red bicycle on a white background",
  size: "1024x1024",
  quality: "medium",
});

const b64 = r.data?.[0]?.b64_json ?? "";
writeFileSync("out.png", Buffer.from(b64, "base64"));
console.log(r.usage); // input/output tokens and cost (USD)
```

## Content-free

Prompts and image bytes are content, and the marketplace never stores content. The prompt
is never logged and never written to the ledger; the base64 bytes and any `revised_prompt`
are relayed to you and dropped. Telemetry keeps the response **byte size** so throughput
stays measurable, and nothing else. See [Data policy](/docs/data-policy).

## Model discovery

```bash
curl -s "https://api.routerplus.com/v1/models?output_modalities=image" \
  -H "Authorization: Bearer $TM_API_KEY"
```

Every row of `GET /v1/models` carries `architecture.output_modalities` — `["text"]`,
`["image"]` or `["video"]` — and the `?output_modalities=` filter takes a comma list of
those values.
The public feed `https://app.routerplus.com/api/models.json` carries `output_modalities`, the full
`supported_parameters` descriptor map for image models, and `max_output_tokens` as the
per-image ceiling. See [GET /v1/models](/docs/api-models) and
[Models & catalog](/docs/models).

## BYOK

Image generation works through an `openai` provider connection. List the exact image ids
you intend to send on the connection — matching is literal, and an id the connection does
not carry is a `404 model_unavailable`. The model's metadata (ceiling, descriptors) still
comes from the house catalog. `usage.cost` is `0`, the marketplace charge is zero, and
OpenAI bills your own account for the render. See [Bring your own key](/docs/byok).

See also: [Models & catalog](/docs/models) · [Pricing & billing](/docs/pricing) ·
[Errors](/docs/errors) · [Wire compatibility](/docs/compat)
