This is the one-page reference for prompting GPT Image 2 — OpenAI's most capable image model, built natively into ChatGPT and available through the OpenAI API as gpt-image-2. It takes no --parameters: a native Thinking Mode reasons about composition, object counts, and lighting before it renders, so you brief it in full descriptive sentences, not tag-soup. Below is the 6-factor formula and copy-paste tables for every modifier that reliably moves the output.
New to it? Start with the best GPT Image 2 prompts roundup, then keep this sheet open while you build your own. For a deeper walk-through, see how to prompt GPT Image 2 for photorealism.
The prompt formula
A strong GPT Image 2 prompt is one descriptive paragraph built from six factors plus a short line of constraints. Cover subject and setting first; the rest sharpen the result. The order is a guide, not a rule — GPT Image 2 reads the whole brief and reasons about it before the first pixel.
- Subject — who or what, described concretely (age, clothing, material, expression).
- Composition / framing — close-up, wide shot, centered, rule-of-thirds, eye-level, low angle.
- Action / pose — what the subject is doing, even if subtle ("mid-stride", "pouring coffee").
- Setting / location — where it happens, with real detail (time of day, weather, surfaces).
- Lighting — the single biggest lever on mood; name a real setup (see the table below).
- Style / medium — photoreal, cinematic, 3D render, watercolor, and so on.
Constraints go on their own line at the end: lens, aspect ratio, resolution, and format — for example, "Shot on 85mm f/1.8. Aspect ratio 3:2. Render in 2K."
Skeleton to copy:
[Subject, described concretely], [composition/framing], [action],
in [setting/location with real detail], lit with [lighting setup],
in the style of [style/medium].
Shot on [lens]. Aspect ratio [X:Y]. Render in [2K/4K].The same skeleton, filled in:
A weathered fisherman in his sixties with a knitted grey sweater, tight
head-and-shoulders portrait, looking just off-camera, on a misty harbour
wall at dawn, lit with soft golden-hour backlighting and gentle rim light,
photorealistic editorial style with natural skin texture.
Shot on 85mm f/1.8 with shallow depth of field. Aspect ratio 4:3. Render in 2K.Why it works: every factor is present and specific, the constraints sit on one clean line, and it reads as a sentence a photographer could shoot — exactly what Thinking Mode is built to reward.
Aspect ratios
State the ratio in plain language ("Aspect ratio 16:9"); there are no flags. GPT Image 2 spans 3:1 ultra-wide to 1:3 tall vertical, and every edge is a multiple of 16. Pick the shape that matches where the image will live.
| Aspect ratio | Shape | Best use |
|---|---|---|
| 3:1 | Ultra-wide | Website hero banners, email headers, panoramic scenes |
| 16:9 | Wide | YouTube thumbnails, slide backgrounds, desktop wallpaper |
| 4:3 | Landscape | Presentation slides, classic monitors, blog headers |
| 1:1 | Square | Profile pics, album art, Instagram grid, product tiles |
| 3:4 | Portrait | Portraits, mobile-first cards, catalogue shots |
| 9:16 | Tall portrait | Stories, Reels, TikTok, phone wallpapers |
| 1:3 | Ultra-tall | Vertical banners, bookmarks, tall-format ads |
Resolution & format
GPT Image 2 renders natively up to 2K (about 2560×1440), roughly double GPT Image 1.5, with optional 4K upscaling. Ask for the resolution and format you need; in the API you set output_format and a quality of low, medium, or high.
| Setting | What it does |
|---|---|
| Render in 1K | Fast drafts and iteration — light and quick to reroll |
| Render in 2K | Native max, the safe default for web, social, and screens |
| Render in 4K | Upscaled detail for print, wallpaper, and tight crops |
| PNG | Lossless; the format to use when you need a transparent alpha channel |
| WebP | Smaller lossless/lossy files, also supports transparency |
| JPEG | Small opaque photos — falls back to a filled background, no alpha |
| Transparent alpha | "Isolated on a transparent background, PNG, no background fill" — true cut-out |
Transparency note: in the OpenAI API set background:"transparent" with output_format:"png" or "webp". JPEG silently produces an opaque background, so never use it for logos or cut-outs.
Lighting modifiers
Lighting sets the mood faster than any other factor. Name a real setup and GPT Image 2 renders it convincingly. Drop one phrase into the lighting slot of the formula.
| Modifier | Effect |
|---|---|
| Three-point softbox setup | Clean, even, flattering studio light — the safe default for portraits and products |
| Beauty dish with reflector fill | Crisp, glamorous front light with soft shadow under the chin — cosmetics and close-ups |
| Golden-hour backlight with long shadows | Warm glow, halo around edges, dreamy and cinematic |
| Soft overcast diffusion | Flat, shadowless, gentle light — natural skin, honest product colour |
| Rim light | Bright edge that separates the subject from a dark background |
| Rembrandt lighting | Small triangle of light on the shadowed cheek — classic portrait look |
| High-key lighting | Bright, airy, low-shadow — beauty, e-commerce, clean and optimistic |
| Low-key lighting | Mostly dark with selective highlights — noir, luxury, tension |
| Neon / practical lights | Coloured urban glow (magenta, cyan) reflecting on wet surfaces |
| Hard direct sunlight | Crisp, high-contrast shadows — bold fashion and street work |
| Candlelight / firelight | Warm, flickering, low and intimate — cosy and romantic scenes |
Lens & camera modifiers
Real lens language tells GPT Image 2 how the scene should compress, blur, and frame. Put these in the constraints line ("Shot on 85mm f/1.8").
| Modifier | Effect |
|---|---|
| 85mm f/1.8 | Flattering portrait compression with creamy background blur |
| 35mm | Natural, documentary, street perspective close to how the eye sees |
| 24mm deep focus | Expansive scenes, interiors, landscapes sharp front to back |
| 100mm macro | Extreme close-up detail — jewellery, food, textures, insects |
| Shallow depth of field | Sharp subject, soft background — isolates and draws the eye |
| Deep depth of field | Everything sharp front to back — landscapes, architecture |
| Bokeh | Soft round out-of-focus highlights behind the subject |
| Tilt-shift | Selective focus band that makes real scenes look miniature |
| Drone / aerial shot | Top-down or high overhead view for scale and pattern |
| Low angle / worm's-eye | Looking up — makes the subject tower and feel heroic |
| Fisheye | Strong curved distortion for a playful, immersive look |
Style & medium modifiers
The style factor decides whether you get a photo, a render, or an illustration. Name one clearly; mixing three fights the model.
| Modifier | Effect |
|---|---|
| Photorealistic | Looks like a real photograph — natural texture, real lighting physics |
| Cinematic | Film colour grade, wide-frame drama, moody contrast and atmosphere |
| 3D render (Octane / Blender) | Clean CGI with glossy materials and studio reflections |
| Watercolour | Soft washes, bleeding edges, visible paper texture |
| Isometric | Angled 3D-tile look for scenes, rooms, and game-style diagrams |
| Flat vector illustration | Bold shapes, limited palette, no gradients — icons and web art |
| Film grain, 35mm analog | Grainy, slightly faded, nostalgic film-camera aesthetic |
| Product studio | Seamless backdrop, crisp reflections, e-commerce-ready lighting |
| Anime / manga | Cel-shaded characters with expressive lines and flat colour |
| Line art / sketch | Pen or pencil linework, minimal or no colour |
| Oil painting | Visible brushstrokes, rich impasto texture, classical feel |
Text-rendering rules
Text rendering is GPT Image 2's headline strength — about 99% English accuracy plus strong Chinese, Japanese, Korean, Hindi, Bengali, and Arabic — so it usually spells clean on the first pass. Follow these rules and it stays crisp even on dense packaging or UI. For ready-made poster skeletons, see the GPT Image 2 prompt templates.
| Rule | Why |
|---|---|
| Wrap the exact words in "quotes" | Tells the model precisely which characters to render — no paraphrasing |
| Keep each string 1–5 words | Short strings render cleanly; long paragraphs risk drift |
| ALL-CAPS, 1–4 words | The most reliable range for crisp, legible headline text |
| Name the font and weight | "bold condensed sans-serif" or "elegant serif" guides the letterforms |
| Put text instruction near the start | The model plans layout early, so the text drives composition |
| Split separate strings | Quote each line on its own so headline and subhead don't merge |
Text example prompt:
A minimalist gym poster with the bold ALL-CAPS headline "NO EXCUSES" in a
heavy condensed sans-serif, centered near the top. Below it, a lone runner
in silhouette on an empty track at dawn, backlit by golden-hour sun with
long shadows. High-contrast, energetic, editorial poster style.
Aspect ratio 3:4. Render in 2K.Why it works: the exact words are quoted, kept to two ALL-CAPS words, the font is named, and the text instruction leads — so the model lays out the type first and builds the image around it.
Editing & consistency
If an image comes back about 80 percent right, don't regenerate — you will lose the parts you liked. Instead describe only the one change and spell out what must stay identical. GPT Image 2 follows instructions closely and holds faces and products consistent, so the rest stays put.
| Phrase | What it does |
|---|---|
| "Change only [X], keep everything else identical" | Isolates the edit to one object or attribute |
| "Keep the pose, lighting, and background exactly the same" | Locks the scene so it stays an edit, not a reroll |
| "Preserve the original aspect ratio and resolution" | Stops the output from resizing or recropping |
| "Keep the person's identity and facial features identical" | Protects a face during relights and background swaps |
| "Match the new background's light direction to the subject" | Makes composites and swaps read as physically real |
| "Use image 1 for the product, image 2 for the face" | Labels reference inputs so each subject stays consistent |
Using the attached photo, change only the jacket to bright red leather.
Keep the pose, lighting, background, and the person's identity exactly the
same. Preserve the original aspect ratio and resolution.Best for: iterating fast — swap a colour, remove an object, or relight a face without rerolling the whole scene. Change one thing per message and check the result before the next tweak.
Copy-paste example prompts
Six complete prompts that assemble the modifiers above. Paste any into ChatGPT or the API and tweak the bracketed parts. For more, browse the best GPT Image 2 prompts.
1. Product hero shot
A matte-black ceramic coffee mug on a wet slate surface, centered
three-quarter view with steam rising, in a dim minimalist studio, lit with
a three-point softbox setup and a subtle rim light, photorealistic product
studio style with crisp reflections and a seamless charcoal backdrop.
Shot on 100mm macro with shallow depth of field. Aspect ratio 1:1. Render in 4K.Best for: e-commerce listings and ads — clean studio light and a macro lens make the material read as premium.
2. Cinematic ultra-wide environment
A lone figure in a long coat standing at a rain-slicked neon crossroads in
a futuristic Tokyo alley at night, wide establishing shot from a low angle,
mid-stride, lit by magenta and cyan neon reflecting off wet asphalt,
cinematic film-grade colour with atmospheric haze.
Shot on 35mm, deep depth of field. Aspect ratio 3:1. Render in 4K.Best for: hero banners and story frames — the 3:1 ultra-wide ratio and neon lighting do the cinematic heavy lifting.
3. Editorial portrait
A confident woman in her thirties with short curls and a tailored linen
blazer, tight head-and-shoulders portrait, looking directly at camera, in a
sunlit loft with a soft blurred window behind her, lit with Rembrandt
lighting and soft overcast fill, photorealistic editorial style with natural
skin texture. Shot on 85mm f/1.8, shallow depth of field. Aspect ratio 3:4.
Render in 2K.Best for: LinkedIn headshots and magazine-style portraits. See the photorealism guide for skin and lighting deep-dives.
4. Transparent-background icon
A single glossy 3D app icon of a folded paper plane in soft indigo-to-violet
gradient, centered, gentle studio reflections and a subtle inner glow,
isolated on a transparent background, PNG with a true alpha channel, no
background fill. Clean modern render.
Aspect ratio 1:1. Render in 2K.Best for: logos, stickers, and UI assets — the transparent-PNG phrasing gives a clean cut-out you can drop onto any surface.
5. Text poster
A retro travel poster with the ALL-CAPS headline "VISIT MARS" in a bold
geometric sans-serif across the top, and a smaller line "EST. 2071" beneath
it. Below the type, a stylised red desert landscape with a distant domed
colony under a pink sky, mid-century illustration style with a warm limited
palette. Aspect ratio 3:4. Render in 2K.Why it works: two short quoted strings, a named font, and the text leading the prompt let GPT Image 2's text engine render both lines crisply.
6. Flat vector web hero
A friendly flat vector illustration of a person watering a large houseplant
in a cozy apartment, centered composition, warm limited palette of terracotta,
sage, and cream, bold clean shapes with no gradients, soft even lighting,
modern flat illustration style for a web hero.
Aspect ratio 16:9. Render in 2K.Best for: landing-page graphics and blog headers where a clean, on-brand illustration beats a photo.
Frequently Asked Questions
How do I set the aspect ratio in GPT Image 2?
State it in plain language inside the prompt — for example, add the line "Aspect ratio 16:9" near the end. GPT Image 2 supports everything from 3:1 ultra-wide to 1:3 tall vertical, including 1:1, 4:3, 3:4, 16:9, and 9:16. Image edges are always multiples of 16. There are no --ar style flags; you describe the ratio in words, which works in ChatGPT and through the OpenAI API alike.
Does GPT Image 2 use --parameters like Midjourney?
No. GPT Image 2 has no --ar, --style, or --v flags. It runs a native Thinking Mode that reasons about composition, object counts, and lighting before rendering, so it reads full descriptive sentences. Write requirements in plain English — "Aspect ratio 16:9", "shot on 85mm f/1.8", "render in 4K" — and brief it like a creative director rather than stacking tags.
What resolution can GPT Image 2 output?
GPT Image 2 renders natively up to 2K, roughly 2560×1440, about double GPT Image 1.5, with optional 4K upscaling. Ask for "2K" for most screen and social work, or "4K" when you need print detail or a wallpaper crop. In the API you also set output_format (PNG, WebP, or JPEG) and a quality of low, medium, or high.
How do I get sharp, correctly spelled text in GPT Image 2?
Wrap the exact words in quotation marks, keep each string short — one to five words, with ALL-CAPS one-to-four-word headlines the most reliable — name the font and weight, and put the text instruction near the start of the prompt. Text rendering is GPT Image 2's headline strength, with about 99% English accuracy and strong results in Chinese, Japanese, Korean, Hindi, Bengali, and Arabic, so it usually renders clean on the first try.
How do I get a transparent background from GPT Image 2?
Ask for it in words: "isolated on a transparent background, PNG with a true alpha channel, no background fill." GPT Image 2 outputs PNG and WebP with a real alpha channel. In the OpenAI API, set background to "transparent" with output_format "png" or "webp" — JPEG silently falls back to an opaque background, so avoid it for cut-outs.
How does editing work in GPT Image 2?
It edits conversationally. Attach the image, describe only the one change you want, and list everything that must stay identical — pose, lighting, background, identity — while preserving the original aspect ratio and resolution. GPT Image 2 follows instructions closely and holds faces and products consistent, so background swaps, retouching, object removal, relighting, and restyling stay clean without redrawing the whole scene.