> Generate text, images, video, audio, realtime voice, and embeddings with a single API. OpenAI-compatible — use any OpenAI SDK by changing the base URL. **Base URL:** `https://gen.pollinations.ai` **Get your API key:** [enter.pollinations.ai](https://enter.pollinations.ai/keys) **Model catalog migration:** model IDs now use `publisher/model` names. Existing aliases remain valid in requests; match catalog entries against both their canonical ID and `aliases` when restoring saved selections. Catalog metadata uses `publisher` (for example, `OpenAI`) instead of `brand`; update clients reading that field. `publisher` identifies the model publisher, not the inference provider. The existing `brand_url` logo field is unchanged. See the [live model catalog](https://gen.pollinations.ai/models) for current IDs and aliases. **Integrations:** [Connect User Wallets](/docs#tag/connect-user-wallets) · [Publish a Model](/docs#tag/publish-a-model) · [Publish an Agent](/docs#tag/publish-an-agent) · [MCP Servers](/docs#tag/mcp-servers) · [CLI](/docs#tag/cli) ## Quick Start ### Text (Python, OpenAI SDK) ```python from openai import OpenAI client = OpenAI(base_url="https://gen.pollinations.ai/v1", api_key="YOUR_API_KEY") response = client.chat.completions.create(model="openai/gpt-5.4-nano", messages=[{"role": "user", "content": "Hello!"}]) print(response.choices[0].message.content) ``` ### Image (URL — no code needed) ```plaintext https://gen.pollinations.ai/image/a%20cat%20in%20space?model=flux ``` ### Audio (cURL) ```bash curl "https://gen.pollinations.ai/audio/Hello%20world?voice=nova" \ -H "Authorization: Bearer YOUR_API_KEY" -o speech.mp3 ``` ### 3D (cURL) ```bash curl "https://gen.pollinations.ai/3d/no_prompt_for_trellis_needed?image=https://inferenceport.ai/img/trellis.jpg&model=microsoft%2Ftrellis-2&resolution=low" \ -H "Authorization: Bearer YOUR_API_KEY" -o model.glb ``` ### Embeddings (OpenAI-compatible) ```bash curl https://gen.pollinations.ai/v1/embeddings \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"openai/text-embedding-3-small","input":"Hello world","dimensions":512}' ``` See `GET /v1/models` for every text, image, audio, video, and embedding model available. ## Authentication All generation requests require an API key from [enter.pollinations.ai](https://enter.pollinations.ai/keys). Model listing endpoints work without authentication. | Type | Prefix | Use case | Rate limits | Description | |------|--------|----------|-------------|-------------| | Secret | `sk_` | Server-side apps | None | Personal developer key. Never expose in client-side code. | | App Key (Connect User Wallets) | `pk_` with redirect URIs | Client apps via OAuth / device flow | None on the App Key itself | Publishable App Key used as the OAuth `client_id`. Users authorize; your app receives a scoped `sk_`. | | Raw publishable | `pk_` with no app binding | Legacy direct spend | 1 pollen / IP / hour | Retained for existing integrations. Do not mint new ones. | > **Note:** Raw publishable keys (`pk_` used as a generation key in browsers) are **legacy**, not beta. New frontend and mobile apps should use **Connect User Wallets**, also called BYOP (Bring Your Own Pollen): register an App Key at [enter.pollinations.ai/keys](https://enter.pollinations.ai/keys), then run the OAuth authorization-code flow with PKCE (or the device flow) to obtain a temporary user-authorized secret key (`sk_`). The legacy fragment redirect and device flow remain supported. Two ways to authenticate generation requests: - Header: `Authorization: Bearer YOUR_API_KEY` - Query param: `?key=YOUR_API_KEY` For detailed integration guidance on user-pays authorization, including OAuth discovery and token exchange, see [Connect User Wallets](https://github.com/pollinations/pollinations/blob/main/BRING_YOUR_OWN_POLLEN.md). ## Text Generation Generate text using OpenAI-compatible Chat Completions and stateless Responses APIs — use an OpenAI SDK by changing the base URL. | Endpoint | Best for | |----------|----------| | `POST /v1/chat/completions` | Full OpenAI compatibility — streaming, tools, vision, structured outputs | | `POST /v1/responses` | Stateless Responses input/output items, semantic streaming events, and function tools | | `POST /v1/messages` | Anthropic Messages API — Claude Code and the Anthropic SDKs | | `GET /text/{prompt}` | Quick prototyping — simple GET, returns plain text | **Available models:** openai/gpt-5.4-nano, openai/gpt-5-nano, openai/gpt-oss-20b, openai/gpt-4o-mini, openai/gpt-5.3-codex, openai/gpt-5.4, openai/gpt-5.4-mini, openai/gpt-5.5, openai/gpt-5.6-sol, openai/gpt-5.6-terra, openai/gpt-5.6-luna, openai/gpt-6-astra, openai/gpt-6-sol, openai/gpt-6.1-sol, openai/gpt-6-luna, inception/mercury-2, inception/mercury-2.5-preview, cohere/command-a-plus, qwen/qwen3-coder-30b-a3b-instruct, mistralai/mistral-small-3.2, mistralai/mistral-small-4, openai/gpt-audio-mini, openai/gpt-audio-1.5, google/gemini-3-flash-preview, google/gemini-3.7-flash, google/gemini-3.8-flash, google/gemini-3.5-flash-lite, google/gemini-2.5-flash-lite, deepseek/deepseek-v4-flash, deepseek/deepseek-v4.1-flash, deepseek/deepseek-v4-flash-vision-exp, google/gemma-4-26b-a4b-it, google/gemma-4-31b-it, deepseek/deepseek-v4-pro, x-ai/grok-4.20, x-ai/grok-4.3, x-ai/grok-4.6, x-ai/grok-4.7, google/gemini-2.5-flash-lite:search, respan/span-01-lite, typesafe/jev-1.13, jaredpalmer/kev-4b, pollinations/midijourney, pollinations/midijourney-large, anthropic/claude-haiku-4.5, anthropic/claude-sonnet-4.6, anthropic/claude-sonnet-5, anthropic/claude-sonnet-5.5, anthropic/claude-opus-4.6, anthropic/claude-opus-4.7, anthropic/claude-opus-5, anthropic/claude-opus-5.5, anthropic/claude-fable-5, anthropic/claude-fable-5.1, perplexity/sonar, moonshotai/kimi-k2.6, moonshotai/kimi-k2.7-code, moonshotai/kimi-k3, poolside/laguna-s-2.1, tencent/hy4-preview, tencent/hy3, inclusionai/ling-3.0-flash-vl, meituan/longcat-2.0, thinkingmachines/inkling-small, thinkingmachines/inkling, nvidia/nemotron-3-ultra, nvidia/nemotron-3.5-lightning, xiaomi/mimo-v2.5, xiaomi/mimo-v2.5-pro, xiaomi/mimo-v2.6-flash, xiaomi/mimo-v2.6-pro, google/gemini-3.1-pro-preview, amazon/nova-micro-v1, amazon/nova-2-lite-v1, z-ai/glm-5.2, z-ai/glm-5.3, z-ai/glm-5.3-flash, z-ai/glm-5.3-flashx, meta/llama-3.3-70b-instruct, meta/llama-4-maverick, meta/llama-4-scout, minimax/minimax-m2.7, minimax/minimax-m3, meta/muse-glimmer-30b, meta/muse-spark-1.2, mistralai/mistral-large-3, qwen/qwen3-coder-next, qwen/qwen3.7-plus, qwen/qwen3.7-max, qwen/qwen3.8-2.4t-a95b, qwen/qwen3.8-27b, qwen/qwen3.8-max, qwen/qwen3.8-max-0902, qwen/qwen3.8-flash, qwen/qwen3.7-flash, qwen/qwen3-vl-30b-a3b-instruct, qwen/qwen3-vl-235b-a22b-thinking, stepfun/step-3.7-flash, stepfun/step-3.5-flash, qwen/qwen3guard-gen-8b ### Responses API Use `supported_endpoints` from [`GET /v1/models`](/v1/models) or [`GET /text/models`](/text/models) to find models that advertise `/v1/responses`. This includes configured built-in providers, community text models with an exact Responses URL, external endpoint agents with an exact Responses URL, and managed prompt agents. ```bash curl https://gen.pollinations.ai/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $POLLINATIONS_API_KEY" \ -d '{ "model": "openai", "input": "Explain why the sky is blue in two sentences.", "store": false }' ``` The endpoint is deliberately stateless. `store` must be `false`; `previous_response_id`, `conversation`, and `prompt` must be null or omitted; `background` must be false or omitted; and encrypted content or reusable item references are rejected. Streaming uses Responses event names and terminal usage events. Direct models preserve the provider's terminal marker; managed-agent streams add one `data: [DONE]` marker. For text models, missing or malformed usage on a completed or incomplete response fails the request. Failed responses may report null usage. These failed requests are not billed, but completed child model calls and charged MCP operations within an agent run remain billable; the outer agent request adds no charge. The stateless surface follows the OpenAI Responses API and OpenResponses item/event vocabulary. It does not claim full OpenResponses conformance: persisted continuation, conversations, compaction, background jobs, Responses WebSocket transport, and normalization of every direct provider stream are outside this subset. Community text models and endpoint agents declare one upstream API and one exact URL. A Responses registration accepts both public APIs: Responses requests use the selected endpoint directly, while Chat Completions requests use the shared stateless adapter. A Chat Completions registration accepts Chat Completions only. Built-in models can have separate routes for the two public APIs; advertising Responses does not mean their Chat requests use the adapter. Managed prompt agents run configured MCP tools on the server. Send previous response items back to continue a conversation; completed tools are not run again. Managed prompt agents accept `reasoning.effort` (Responses) and `reasoning_effort` (Chat Completions). Reasoning summaries are not supported: a non-null `reasoning.summary` returns HTTP 400. ### Anthropic Messages API Models that list `/v1/messages` in `supported_endpoints` — every text model that supports Chat Completions — also accept Anthropic Messages requests. Point Claude Code or an Anthropic SDK at `https://gen.pollinations.ai` and authenticate with a bearer token: ```bash export ANTHROPIC_BASE_URL=https://gen.pollinations.ai export ANTHROPIC_AUTH_TOKEN=$POLLINATIONS_API_KEY export ANTHROPIC_MODEL=openai claude ``` ```python import os import anthropic client = anthropic.Anthropic( base_url="https://gen.pollinations.ai", auth_token=os.environ["POLLINATIONS_API_KEY"], ) message = client.messages.create( model="openai", max_tokens=1024, messages=[{"role": "user", "content": "Hello"}], ) ``` ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://gen.pollinations.ai", authToken: process.env.POLLINATIONS_API_KEY, }); const message = await client.messages.create({ model: "openai", max_tokens: 1024, messages: [{ role: "user", content: "Hello" }], }); ``` Requests run as Chat Completions requests: balance checks, key permissions, rate limits, caching and billing are the same. Streaming, tools, images, system prompts and stop sequences depend on the selected model's capabilities; see [`/text/models`](/text/models). `cache_control` uses the same provider support as Chat Completions (see Prompt caching below); custom cache TTLs are not supported. `thinking` sets `reasoning_effort` (`output_config.effort` for adaptive thinking), and provider reasoning returns as `thinking` blocks. Usage reports `input_tokens`, `output_tokens`, `cache_read_input_tokens` and `cache_creation_input_tokens`; a response without provider usage fails, and a stream ends with an `error` event. Errors use Anthropic's error shape. `/v1/messages/count_tokens`, batches, files, server tools and `x-api-key` authentication are not supported. Claude Code sends `cache_control` automatically. Fireworks-hosted models that reject this field cannot currently be used with Claude Code; see [the compatibility issue](https://github.com/pollinations/pollinations/issues/15682). ### Media models in conversations Image, video, audio and 3D models that advertise these endpoints in [`/models`](/models) accept a text prompt. Only the last user message is used; history, instructions and text-generation settings are ignored. Its text parts (or a string Responses `input`) form the prompt. Image parts (`image_url` in Chat, `input_image` in Responses, as URLs or data URIs) are the source images of image models and the start frame of video models that list `image` under `input_modalities`, exactly as `/v1/images/edits` does; other models, including 3D, return HTTP 400 for them, and any other attachment type returns HTTP 400. Use the native media endpoints for generation settings. Text models with `video` under `input_modalities` (for example `inclusionai/ling-3.0-flash-vl`) accept `video_url` parts the same way they accept `image_url`: a public `https://` URL or a `data:video/...;base64,...` data URI (Gemini models also accept YouTube and `gs://` URLs). Video usage is metered from the provider's reported `video_tokens` detail and billed against the model's video prompt rate. Empty prompts, malformed Unicode and prompts consisting only of `.` or `..` return HTTP 400. Reference-required models return their normal missing-input error. Dialogue models expect one `: ` turn per line, just like `/audio`. Community speech models available only through `/v1/audio/speech` are not included. ```bash curl https://gen.pollinations.ai/v1/responses \ -H "Authorization: Bearer $POLLINATIONS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"flux","input":"A lighthouse at dawn"}' curl https://gen.pollinations.ai/v1/chat/completions \ -H "Authorization: Bearer $POLLINATIONS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"flux","messages":[{"role":"user","content":"A lighthouse at dawn"}]}' ``` Both return assistant text: a Markdown image embed for images, or a Markdown link for audio, video and 3D, followed by the plain public file URL. The URL is also in the `Link` header. With `stream: true`, events arrive after generation finishes. Media uses its normal billing units, not text tokens: Responses returns `usage: null`; Chat JSON omits `usage`. Chat streaming chunks contain `usage: null`, with no final usage chunk. Video uses the native model or provider's default duration. ### Reasoning Use `reasoning_effort` to control reasoning on models that advertise reasoning support. ```bash # POST /v1/chat/completions — OpenAI-compatible response curl https://gen.pollinations.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $POLLINATIONS_API_KEY" \ -d '{ "model": "openai", "reasoning_effort": "high", "messages": [ { "role": "user", "content": "Prove that there are infinitely many prime numbers." } ] }' ``` ```bash # POST /text — plain-text response curl https://gen.pollinations.ai/text \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $POLLINATIONS_API_KEY" \ -d '{ "model": "openai", "reasoning_effort": "medium", "messages": [ { "role": "user", "content": "Design a URL shortener. Outline the key tradeoffs." } ] }' ``` ### Prompt caching On Gemini, Claude, and Nova models, a large static prompt prefix can be cached so repeat requests bill it at a fraction of the input rate. Mark the end of the static prefix with `cache_control` on a content block (not on the message); everything before the marker must be byte-identical across requests, everything dynamic goes after. The first request creates the cache (`usage` reports `cache_creation_input_tokens`); repeat requests within the TTL report `prompt_tokens_details.cached_tokens` at the discounted rate. ```json { "model": "google/gemini-2.5-flash-lite", "messages": [ { "role": "system", "content": [ { "type": "text", "text": "", "cache_control": { "type": "ephemeral" } } ] }, { "role": "user", "content": "" } ] } ``` **Gemini** — the prefix must be at least ~2,048 tokens (~4,096 on Gemini 3 models). Requests with tools are not cached — including built-in tools, so `google/gemini-3.7-flash`, `google/gemini-3-flash-preview`, `google/gemini-3.1-pro-preview`, and the search variants only cache when tools are disabled (`"tools": []`) or a JSON `response_format` is set; `google/gemini-2.5-flash-lite` and `google/gemini-3.5-flash-lite` cache by default. Cache creates bill at the standard input rate plus a storage fee for the 1-hour TTL ($1 per 1M cached tokens on Flash models, $4.50 on Pro); hits bill at ~10% of input. The storage fee means caching pays off only when the prefix is reused often — roughly a dozen reuses per hour on the cheapest models. **Claude** — all Claude models cache. The prefix minimum varies by model: 512 tokens on `anthropic/claude-fable-5`, `anthropic/claude-fable-5.1`, and `anthropic/claude-opus-5`, and 1,024 on `anthropic/claude-sonnet-4.6`; other models have higher minimums. Tool definitions are cacheable. `anthropic/claude-fable-5.1` accepts only automatic or disabled tool choice; forcing any or a named tool returns a 400. Cache creates bill at 1.25× the input rate (no storage fee); hits bill at 10% of input, or 2.5% on `anthropic/claude-fable-5.1`. The cache lives ~5 minutes, refreshed on each hit. **Nova** — `nova` and `nova-fast` cache. The prefix must be at least ~1,000 tokens (up to 20K tokens cacheable). Cache creates are free; hits bill at 25% of input. ~5-minute TTL. Models that advertise `/v1/responses` also accept OpenAI's cache controls. Set `prompt_cache_options.mode` to `explicit` and place `prompt_cache_breakpoint: { "mode": "explicit" }` on the content block ending each stable prefix (up to four). Chat requests adapted to Responses preserve these markers; the existing `cache_control: { "type": "ephemeral" }` marker is translated to the same explicit breakpoint. Managed prompt agents apply an explicit request without caller markers to their configured static prompt. ### Typed decisions (`typesafe/jev-1.13`) `typesafe/jev-1.13` (aliases `jev` and `typesafe/jev`) returns calibrated judgments instead of free text. Post `state` and a map of `questions` to `POST /alpha/decisions`; each question is a `choice`, `score`, or `noul`, and each is answered independently under the key you supplied. `model` defaults to `jev`. ```json { "state": "My payouts have been failing for 3 days.", "questions": { "department": { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "Payment issues", "technical": "Product failures" } }, "is_urgent": { "type": "noul", "instructions": "Does this convey urgency?" } } } ``` The response carries `answers`, one field per question, each with `type` and its native fields (`choice` + `confidence` + `probabilities`, `score` + `legend` + `confidence` + `probabilities`, or `noul`), plus `usage` with `input_tokens` and `output_tokens`. See the [TypeSafe API reference](https://docs.typesafe.ai/api) for the native request and answer shapes. ```json { "id": "dec-…", "model": "typesafe/jev-1.13", "provider": "TypeSafe", "answers": { "department": { "type": "choice", "choice": "billing", "confidence": 0.82, "probabilities": { "billing": 0.91, "technical": 0.09 } }, "is_urgent": { "type": "noul", "noul": 0.87 } }, "usage": { "input_tokens": 312, "output_tokens": 48 } } ``` `state`, `instructions`, and criteria values accept a string or arbitrary JSON. There is no streaming; the answers arrive in one response. The same model is also reachable from an OpenAI client on `/v1/chat/completions`: put the identical request JSON in the last `user` message as a string, and the answers come back as `message.content`. Earlier turns, system instructions, and text-generation settings are ignored. With `stream: true` the finished answers arrive as one content chunk followed by the usage chunk. Prefer `/alpha/decisions` where you can post the native shape. Supply relevant facts in `state`; Jev can be confident even when facts are missing. Interpret scores using `legend`, and handle counting, arithmetic, and date comparisons in code. Questions are evaluated independently. The context limit is 64k tokens for `state` and all questions together, and 32k for `state` plus the longest question. ## Image Generation Generate images from text prompts via a simple GET request. Returns JPEG, PNG, or SVG depending on the selected model. ``` https://gen.pollinations.ai/image/a%20cat%20in%20space?model=flux ``` **Available models:** krea/krea-2-medium, lykon/dreamshaper-8-lcm, black-forest-labs/flux.1-kontext-pro, black-forest-labs/flux.1.1-pro, black-forest-labs/flux.2-pro, black-forest-labs/flux.2-flex, black-forest-labs/flux.2-max, microsoft/mai-image-2.6-flash, microsoft/mai-image-2.6, google/gemini-2.5-flash-image, google/gemini-3.1-flash-image, google/gemini-3.1-flash-lite-image, google/gemini-3-pro-image, bytedance/seedream-5.0-lite, bytedance/seedream-5.0-pro, bytedance/seedream-4.0, bytedance/seedream-4.5, ideogram-ai/ideogram-v4-turbo, ideogram-ai/ideogram-v4-balanced, ideogram-ai/ideogram-v4-quality, openai/gpt-image-1-mini, openai/gpt-image-1.5, openai/gpt-image-2, openai/gpt-image-2.5-flare, openai/gpt-image-2.5-sunburst, black-forest-labs/flux.1-schnell, tongyi-mai/z-image-turbo, alibaba/wan-2.7-image, alibaba/wan-2.7-image-pro, qwen/qwen-image, qwen/qwen-image-2.1, qwen/qwen-image-3, x-ai/grok-imagine-image, x-ai/grok-imagine-image-quality, x-ai/grok-imagine-image-2.0, recraft/recraft-v4.1-vector, recraft/recraft-v4.1-flash, black-forest-labs/flux.2-klein-4b, prunaai/p-image, prunaai/p-image-edit, inferenceport-ai/lightning-image-turbo ### Community image models Community image models use a `community/owner/model` id and support generation through `/image/{prompt}` and `/v1/images/generations`. The registration test adds image input and `/v1/images/edits` metadata when the registrant's edit endpoint succeeds. OpenAI-compatible responses default to `b64_json`; set `response_format: "url"` for a stored media URL. See `/image/models` for the live model list and supported endpoints. ## Video Generation Generate videos from text prompts or reference images. Returns MP4. ``` https://gen.pollinations.ai/video/sunset%20timelapse?model=veo&duration=4 ``` **Available models:** google/veo-3.1-fast, google/gemini-omni-1.1-flash, bytedance/seedance-1-pro-fast, bytedance/seedance-2.0, bytedance/seedance-2.0-mini, bytedance/seedance-2.0-fast, alibaba/wan-2.6, alibaba/wan-2.2-fast, alibaba/wan-2.7, alibaba/wan-3.0, x-ai/grok-imagine-video, x-ai/grok-imagine-video-1.5, bytedance/seedance-2.5, alibaba/happyhorse-1.1, heygen/heygen-video-1, minimax/minimax-h3, minimax/minimax-h3-max, minimax/minimax-h3-max-turbo, prunaai/p-video ### Community video models Community video models use a `community/owner/model` id and work on `/video/{prompt}`, `/image/{prompt}`, and `/v1/images/generations`. See `/video/models` for the live catalog and [Publish a Model](https://github.com/pollinations/pollinations/blob/main/BRING_YOUR_OWN_MODEL.md) for the synchronous publisher contract. ## Realtime OpenAI-compatible Realtime WebSocket for voice, multimodal, and transcription sessions. | Endpoint | Description | |----------|-------------| | `GET /realtime` | Pollinations Realtime session (`model=openai/gpt-realtime-2.1`) | | `GET /v1/realtime` | WebSocket Realtime session (`model=openai/gpt-realtime-2.1`) | Requires an API key with positive balance. Server clients can use `Authorization: Bearer `; browser WebSocket clients can use `?key=pk_...`. The WebSocket settles one billing event when the session closes. Selecting `elevenlabs/scribe-v2-realtime` creates a transcription session automatically; other realtime models create voice and multimodal sessions. Events sent and received over both routes use the OpenAI Realtime protocol. See OpenAI's [Realtime WebSocket events guide](https://developers.openai.com/api/docs/guides/realtime-websocket#sending-and-receiving-events). ```js import WebSocket from "ws"; // Server: Bearer auth. Browser: append `&key=pk_...` instead (headers aren't settable). const ws = new WebSocket( "wss://gen.pollinations.ai/v1/realtime?model=openai/gpt-realtime-2.1", { headers: { Authorization: `Bearer ${process.env.POLLINATIONS_API_KEY}` } }, ); ws.on("open", () => ws.send(JSON.stringify({ type: "session.update", session: { type: "realtime", instructions: "Be concise." }, }))); ws.on("message", (m) => console.log(JSON.parse(m.toString()))); ``` **Browser audio:** play the model's audio through an `