Image Generator
Generate AI images from text or reference photos. Product shots, character art, posters with readable text, photoreal portraits, high resolution in seconds.
Runs on
-
OpenAI
-
Google
-
xAI
Examples
Cobalt tide in a Paraty studio
Nano Banana 2. A prompt-only ceramic studio scene where cobalt glaze rises from a handmade bowl as a miniature translucent wave.
How to use Image Generator
A compact decision guide for placing this pipeline in a workflow.
Best for
- Niche/specific subjects with no stock footage (historical recreations, sci-fi concepts, fantasy scenes)
- Conceptual visuals (abstract ideas, future technology, impossible scenes)
- Consistent visual style across multiple images (same lighting, color palette)
- Product visualizations and mockups
- Posters with readable text across Latin and non-Latin scripts
Tips
- Be extremely specific: 'Photorealistic futuristic electric vehicle on a mountain road, dramatic low angle, golden hour, cinematic lighting' not 'car on road'
- If the output drifts toward illustration when you want a photo, add 'photorealistic' explicitly
- Style anchors that work: 'cinematic', 'dramatic lighting', 'volumetric fog', 'shallow depth of field'
- Use reference images to keep the subject, pose, or product shape consistent across multiple generations
Image Generator API
Call this pipeline from your own code. One request dispatches a run; the model is an input field, not a separate endpoint.
Get an API token- inputs
- 7
- Pricing
- 0.2 to 6.4 credits
| Input | Type | default | Description |
|---|---|---|---|
prompt | string | default— | Text description of the image. The AI enhances this with composition, lighting, and style. |
reference_images | array | default— | Optional image references. Auto and the Gemini models accept up to 14; GPT Image 2 accepts up to 5. The selected model limit is applied automatically. |
aspect_ratio | enum | default1:1 | Output aspect ratio.1:1 · 16:9 · 9:16 +21:116:99:164:33:4 |
enhance_prompt | boolean | defaulttrue | Improve the prompt for the selected model before generation. Turn off to send your prompt unchanged. |
model | enum | defaultauto | Image engine to use. Leave unset and Pipe2 routes to the best fit for your prompt.auto · gemini-3-1-flash-image · gemini-3-pro-image +3autogemini-3-1-flash-imagegemini-3-pro-imagegpt-image-2grok-imagine-imagegrok-imagine-image-quality |
quality | enum | defaultauto | Engine-aware. Auto picks the balanced tier for the chosen engine — GPT Image 2 medium, Gemini 3.1 Flash 2K. Fast renders smaller and cheaper (GPT Image 2 low, Gemini 3.1 Flash 1K); hd takes the top tier, which on GPT Image 2 costs many times auto. Gemini 3 Pro changes size at one price.auto · fast · hdautofasthd |
reference_videos | array | default— | Optional: one video reference; the first ~8 seconds are sampled. Only Gemini 3.1 Flash supports video-to-image. |
Point any MCP client at pipe2 and your agent gets these tools. It finds this pipeline, reads the same schema above, then runs it.
list_pipelinesget_pipeline_schemarun_pipelineget_pipeline_run_statusrequest_upload
https://mcp.pipe2.ai/mcp {
"mcpServers": {
"pipe2ai": {
"url": "https://mcp.pipe2.ai/mcp",
"headers": { "Authorization": "Bearer YOUR_TOKEN" }
}
}
}Recipes using this pipeline
Show all recipes-
One prompt → a vertical AI dance reel
Generate a labeled dance-move grid, use it as the choreography reference for a music-synced dance, then combine both into a vertical reel.
-
An ordinary box opens onto a tiny impossible world
Create a silent reveal of a tiny moonlit ocean inside a hinged wooden box: build the open scene, close its lid with an image edit, then animate the discovery.
Frequently Asked Questions
What can the image generator make?
Do I need both a prompt and reference images?
How many reference images can I use?
What image quality can I expect?
Can it render text inside the image?
How many credits does it cost?
AI Image Generator
Generate high-quality images from text prompts or reference photos. Product shots, character art, posters with readable text, photoreal portraits, describe what you want and get a finished image in seconds.
What you can make
- Product photography: catalog shots, lifestyle scenes, and aspirational creative without a studio
- Character art: concept design and variations from a description or reference photo
- Posters with readable text: typography rendered cleanly across Latin and non-Latin scripts (Japanese, Korean, Chinese, Hindi, Bengali, Arabic)
- Editorial imagery: cover art, hero images, social posts in any style
- Brand-anchored variations: upload a product photo or reference and generate consistent visuals around it
How it works
- Describe the image: write what you want, or upload reference photos. Auto and the Gemini models accept up to 14; GPT Image 2 accepts up to 5.
- Pick a model: or leave it on Auto and the router picks for you (cheap drafts, balanced HD, premium typography)
- How Auto routes: text-heavy posters route to the model strong on typography; photoreal portraits route to the model strong on faces and product detail
- Refine: change the prompt, add references, or swap the aspect ratio and run again
Prompt tips
- Be specific about subject, style, and lighting, "a vintage brass telescope on a wooden desk, warm afternoon sunlight" beats "telescope"
- For text in the image, quote what you want rendered, "a poster reading 'GRAND OPENING' in bold sans-serif"
- Upload reference images to anchor pose, color, or product shape across multiple generations
- Use aspect ratio to match your output target, 1:1 for social, 16:9 for hero banners, 9:16 for vertical posts