This is the one-page reference for prompting GPT Image 2 — OpenAI's most capable image model, built natively into ChatGPT and available through the OpenAI API as gpt-image-2. It takes no --parameters: a native Thinking Mode reasons about composition, object counts, and lighting before it renders, so you brief it in full descriptive sentences, not tag-soup. Below is the 6-factor formula and copy-paste tables for every modifier that reliably moves the output.

New to it? Start with the best GPT Image 2 prompts roundup, then keep this sheet open while you build your own. For a deeper walk-through, see how to prompt GPT Image 2 for photorealism.

Advertisement

The prompt formula

A strong GPT Image 2 prompt is one descriptive paragraph built from six factors plus a short line of constraints. Cover subject and setting first; the rest sharpen the result. The order is a guide, not a rule — GPT Image 2 reads the whole brief and reasons about it before the first pixel.

  1. Subject — who or what, described concretely (age, clothing, material, expression).
  2. Composition / framing — close-up, wide shot, centered, rule-of-thirds, eye-level, low angle.
  3. Action / pose — what the subject is doing, even if subtle ("mid-stride", "pouring coffee").
  4. Setting / location — where it happens, with real detail (time of day, weather, surfaces).
  5. Lighting — the single biggest lever on mood; name a real setup (see the table below).
  6. Style / medium — photoreal, cinematic, 3D render, watercolor, and so on.

Constraints go on their own line at the end: lens, aspect ratio, resolution, and format — for example, "Shot on 85mm f/1.8. Aspect ratio 3:2. Render in 2K."

Skeleton to copy:

[Subject, described concretely], [composition/framing], [action],
in [setting/location with real detail], lit with [lighting setup],
in the style of [style/medium].
Shot on [lens]. Aspect ratio [X:Y]. Render in [2K/4K].

The same skeleton, filled in:

A weathered fisherman in his sixties with a knitted grey sweater, tight
head-and-shoulders portrait, looking just off-camera, on a misty harbour
wall at dawn, lit with soft golden-hour backlighting and gentle rim light,
photorealistic editorial style with natural skin texture.
Shot on 85mm f/1.8 with shallow depth of field. Aspect ratio 4:3. Render in 2K.

Why it works: every factor is present and specific, the constraints sit on one clean line, and it reads as a sentence a photographer could shoot — exactly what Thinking Mode is built to reward.

Aspect ratios

State the ratio in plain language ("Aspect ratio 16:9"); there are no flags. GPT Image 2 spans 3:1 ultra-wide to 1:3 tall vertical, and every edge is a multiple of 16. Pick the shape that matches where the image will live.

Aspect ratioShapeBest use
3:1Ultra-wideWebsite hero banners, email headers, panoramic scenes
16:9WideYouTube thumbnails, slide backgrounds, desktop wallpaper
4:3LandscapePresentation slides, classic monitors, blog headers
1:1SquareProfile pics, album art, Instagram grid, product tiles
3:4PortraitPortraits, mobile-first cards, catalogue shots
9:16Tall portraitStories, Reels, TikTok, phone wallpapers
1:3Ultra-tallVertical banners, bookmarks, tall-format ads

Resolution & format

GPT Image 2 renders natively up to 2K (about 2560×1440), roughly double GPT Image 1.5, with optional 4K upscaling. Ask for the resolution and format you need; in the API you set output_format and a quality of low, medium, or high.

SettingWhat it does
Render in 1KFast drafts and iteration — light and quick to reroll
Render in 2KNative max, the safe default for web, social, and screens
Render in 4KUpscaled detail for print, wallpaper, and tight crops
PNGLossless; the format to use when you need a transparent alpha channel
WebPSmaller lossless/lossy files, also supports transparency
JPEGSmall opaque photos — falls back to a filled background, no alpha
Transparent alpha"Isolated on a transparent background, PNG, no background fill" — true cut-out

Transparency note: in the OpenAI API set background:"transparent" with output_format:"png" or "webp". JPEG silently produces an opaque background, so never use it for logos or cut-outs.

Lighting modifiers

Lighting sets the mood faster than any other factor. Name a real setup and GPT Image 2 renders it convincingly. Drop one phrase into the lighting slot of the formula.

ModifierEffect
Three-point softbox setupClean, even, flattering studio light — the safe default for portraits and products
Beauty dish with reflector fillCrisp, glamorous front light with soft shadow under the chin — cosmetics and close-ups
Golden-hour backlight with long shadowsWarm glow, halo around edges, dreamy and cinematic
Soft overcast diffusionFlat, shadowless, gentle light — natural skin, honest product colour
Rim lightBright edge that separates the subject from a dark background
Rembrandt lightingSmall triangle of light on the shadowed cheek — classic portrait look
High-key lightingBright, airy, low-shadow — beauty, e-commerce, clean and optimistic
Low-key lightingMostly dark with selective highlights — noir, luxury, tension
Neon / practical lightsColoured urban glow (magenta, cyan) reflecting on wet surfaces
Hard direct sunlightCrisp, high-contrast shadows — bold fashion and street work
Candlelight / firelightWarm, flickering, low and intimate — cosy and romantic scenes

Lens & camera modifiers

Real lens language tells GPT Image 2 how the scene should compress, blur, and frame. Put these in the constraints line ("Shot on 85mm f/1.8").

ModifierEffect
85mm f/1.8Flattering portrait compression with creamy background blur
35mmNatural, documentary, street perspective close to how the eye sees
24mm deep focusExpansive scenes, interiors, landscapes sharp front to back
100mm macroExtreme close-up detail — jewellery, food, textures, insects
Shallow depth of fieldSharp subject, soft background — isolates and draws the eye
Deep depth of fieldEverything sharp front to back — landscapes, architecture
BokehSoft round out-of-focus highlights behind the subject
Tilt-shiftSelective focus band that makes real scenes look miniature
Drone / aerial shotTop-down or high overhead view for scale and pattern
Low angle / worm's-eyeLooking up — makes the subject tower and feel heroic
FisheyeStrong curved distortion for a playful, immersive look
Advertisement

Style & medium modifiers

The style factor decides whether you get a photo, a render, or an illustration. Name one clearly; mixing three fights the model.

ModifierEffect
PhotorealisticLooks like a real photograph — natural texture, real lighting physics
CinematicFilm colour grade, wide-frame drama, moody contrast and atmosphere
3D render (Octane / Blender)Clean CGI with glossy materials and studio reflections
WatercolourSoft washes, bleeding edges, visible paper texture
IsometricAngled 3D-tile look for scenes, rooms, and game-style diagrams
Flat vector illustrationBold shapes, limited palette, no gradients — icons and web art
Film grain, 35mm analogGrainy, slightly faded, nostalgic film-camera aesthetic
Product studioSeamless backdrop, crisp reflections, e-commerce-ready lighting
Anime / mangaCel-shaded characters with expressive lines and flat colour
Line art / sketchPen or pencil linework, minimal or no colour
Oil paintingVisible brushstrokes, rich impasto texture, classical feel

Text-rendering rules

Text rendering is GPT Image 2's headline strength — about 99% English accuracy plus strong Chinese, Japanese, Korean, Hindi, Bengali, and Arabic — so it usually spells clean on the first pass. Follow these rules and it stays crisp even on dense packaging or UI. For ready-made poster skeletons, see the GPT Image 2 prompt templates.

RuleWhy
Wrap the exact words in "quotes"Tells the model precisely which characters to render — no paraphrasing
Keep each string 1–5 wordsShort strings render cleanly; long paragraphs risk drift
ALL-CAPS, 1–4 wordsThe most reliable range for crisp, legible headline text
Name the font and weight"bold condensed sans-serif" or "elegant serif" guides the letterforms
Put text instruction near the startThe model plans layout early, so the text drives composition
Split separate stringsQuote each line on its own so headline and subhead don't merge

Text example prompt:

A minimalist gym poster with the bold ALL-CAPS headline "NO EXCUSES" in a
heavy condensed sans-serif, centered near the top. Below it, a lone runner
in silhouette on an empty track at dawn, backlit by golden-hour sun with
long shadows. High-contrast, energetic, editorial poster style.
Aspect ratio 3:4. Render in 2K.

Why it works: the exact words are quoted, kept to two ALL-CAPS words, the font is named, and the text instruction leads — so the model lays out the type first and builds the image around it.

Editing & consistency

If an image comes back about 80 percent right, don't regenerate — you will lose the parts you liked. Instead describe only the one change and spell out what must stay identical. GPT Image 2 follows instructions closely and holds faces and products consistent, so the rest stays put.

PhraseWhat it does
"Change only [X], keep everything else identical"Isolates the edit to one object or attribute
"Keep the pose, lighting, and background exactly the same"Locks the scene so it stays an edit, not a reroll
"Preserve the original aspect ratio and resolution"Stops the output from resizing or recropping
"Keep the person's identity and facial features identical"Protects a face during relights and background swaps
"Match the new background's light direction to the subject"Makes composites and swaps read as physically real
"Use image 1 for the product, image 2 for the face"Labels reference inputs so each subject stays consistent
Using the attached photo, change only the jacket to bright red leather.
Keep the pose, lighting, background, and the person's identity exactly the
same. Preserve the original aspect ratio and resolution.

Best for: iterating fast — swap a colour, remove an object, or relight a face without rerolling the whole scene. Change one thing per message and check the result before the next tweak.

Copy-paste example prompts

Six complete prompts that assemble the modifiers above. Paste any into ChatGPT or the API and tweak the bracketed parts. For more, browse the best GPT Image 2 prompts.

1. Product hero shot

A matte-black ceramic coffee mug on a wet slate surface, centered
three-quarter view with steam rising, in a dim minimalist studio, lit with
a three-point softbox setup and a subtle rim light, photorealistic product
studio style with crisp reflections and a seamless charcoal backdrop.
Shot on 100mm macro with shallow depth of field. Aspect ratio 1:1. Render in 4K.

Best for: e-commerce listings and ads — clean studio light and a macro lens make the material read as premium.

2. Cinematic ultra-wide environment

A lone figure in a long coat standing at a rain-slicked neon crossroads in
a futuristic Tokyo alley at night, wide establishing shot from a low angle,
mid-stride, lit by magenta and cyan neon reflecting off wet asphalt,
cinematic film-grade colour with atmospheric haze.
Shot on 35mm, deep depth of field. Aspect ratio 3:1. Render in 4K.

Best for: hero banners and story frames — the 3:1 ultra-wide ratio and neon lighting do the cinematic heavy lifting.

3. Editorial portrait

A confident woman in her thirties with short curls and a tailored linen
blazer, tight head-and-shoulders portrait, looking directly at camera, in a
sunlit loft with a soft blurred window behind her, lit with Rembrandt
lighting and soft overcast fill, photorealistic editorial style with natural
skin texture. Shot on 85mm f/1.8, shallow depth of field. Aspect ratio 3:4.
Render in 2K.

Best for: LinkedIn headshots and magazine-style portraits. See the photorealism guide for skin and lighting deep-dives.

4. Transparent-background icon

A single glossy 3D app icon of a folded paper plane in soft indigo-to-violet
gradient, centered, gentle studio reflections and a subtle inner glow,
isolated on a transparent background, PNG with a true alpha channel, no
background fill. Clean modern render.
Aspect ratio 1:1. Render in 2K.

Best for: logos, stickers, and UI assets — the transparent-PNG phrasing gives a clean cut-out you can drop onto any surface.

5. Text poster

A retro travel poster with the ALL-CAPS headline "VISIT MARS" in a bold
geometric sans-serif across the top, and a smaller line "EST. 2071" beneath
it. Below the type, a stylised red desert landscape with a distant domed
colony under a pink sky, mid-century illustration style with a warm limited
palette. Aspect ratio 3:4. Render in 2K.

Why it works: two short quoted strings, a named font, and the text leading the prompt let GPT Image 2's text engine render both lines crisply.

6. Flat vector web hero

A friendly flat vector illustration of a person watering a large houseplant
in a cozy apartment, centered composition, warm limited palette of terracotta,
sage, and cream, bold clean shapes with no gradients, soft even lighting,
modern flat illustration style for a web hero.
Aspect ratio 16:9. Render in 2K.

Best for: landing-page graphics and blog headers where a clean, on-brand illustration beats a photo.

Frequently Asked Questions

How do I set the aspect ratio in GPT Image 2?

State it in plain language inside the prompt — for example, add the line "Aspect ratio 16:9" near the end. GPT Image 2 supports everything from 3:1 ultra-wide to 1:3 tall vertical, including 1:1, 4:3, 3:4, 16:9, and 9:16. Image edges are always multiples of 16. There are no --ar style flags; you describe the ratio in words, which works in ChatGPT and through the OpenAI API alike.

Does GPT Image 2 use --parameters like Midjourney?

No. GPT Image 2 has no --ar, --style, or --v flags. It runs a native Thinking Mode that reasons about composition, object counts, and lighting before rendering, so it reads full descriptive sentences. Write requirements in plain English — "Aspect ratio 16:9", "shot on 85mm f/1.8", "render in 4K" — and brief it like a creative director rather than stacking tags.

What resolution can GPT Image 2 output?

GPT Image 2 renders natively up to 2K, roughly 2560×1440, about double GPT Image 1.5, with optional 4K upscaling. Ask for "2K" for most screen and social work, or "4K" when you need print detail or a wallpaper crop. In the API you also set output_format (PNG, WebP, or JPEG) and a quality of low, medium, or high.

How do I get sharp, correctly spelled text in GPT Image 2?

Wrap the exact words in quotation marks, keep each string short — one to five words, with ALL-CAPS one-to-four-word headlines the most reliable — name the font and weight, and put the text instruction near the start of the prompt. Text rendering is GPT Image 2's headline strength, with about 99% English accuracy and strong results in Chinese, Japanese, Korean, Hindi, Bengali, and Arabic, so it usually renders clean on the first try.

How do I get a transparent background from GPT Image 2?

Ask for it in words: "isolated on a transparent background, PNG with a true alpha channel, no background fill." GPT Image 2 outputs PNG and WebP with a real alpha channel. In the OpenAI API, set background to "transparent" with output_format "png" or "webp" — JPEG silently falls back to an opaque background, so avoid it for cut-outs.

How does editing work in GPT Image 2?

It edits conversationally. Attach the image, describe only the one change you want, and list everything that must stay identical — pose, lighting, background, identity — while preserving the original aspect ratio and resolution. GPT Image 2 follows instructions closely and holds faces and products consistent, so background swaps, retouching, object removal, relighting, and restyling stay clean without redrawing the whole scene.

Advertisement