Photorealism from GPT Image 2 is mostly a prompting problem, not a model problem. This guide gives you a single 6-part formula — subject, composition, action, setting, lighting, and style, plus a short line of constraints — then 10 complete prompts you can paste into ChatGPT right now, plus the handful of mistakes that make images look fake. For the full library, start with the best GPT Image 2 prompts.
Why GPT Image 2 can look like a real photograph
GPT Image 2 is OpenAI's most capable image model, and it runs a Native Thinking Mode that reasons about composition, object counts, lighting, and constraints before it renders the first pixel. That reasoning step is the difference: it works out the relationships in your scene up front, so a complete, specific brief needs far fewer rerolls, and real camera and lighting language — a shallow 85mm depth of field, a three-point softbox, golden-hour backlight — actually lands.
It also renders natively up to 2K (about 2560×1440, roughly double GPT Image 1.5) with optional 4K upscaling, so fine texture like skin pores and fabric weave survives instead of smearing. The upshot: photorealism comes from briefing it like a photographer, not from stacking quality keywords. Describe the scene precisely and it fills in physically plausible detail; feed it tag-soup and it lands on the generic AI average.
The photorealism formula
Every photoreal prompt is one descriptive paragraph built from six factors, in order, followed by a short line of constraints. Write it as full sentences — act like a creative director briefing a shoot — not a comma list. Here is the skeleton:
[SUBJECT: who or what, with specific, concrete detail]
[COMPOSITION & FRAMING: shot size, angle, placement in frame]
[ACTION: what the subject is doing right now]
[SETTING: where they are, with real environmental detail]
[LIGHTING: a named, real lighting setup and its direction]
[STYLE: the photographic look — editorial, documentary, film stock]
[TEXTURE CUES: physical details that read as real — skin pores, fabric weave, condensation]
Shot on [lens, e.g. 85mm f/1.8]. Aspect ratio 3:2, 4K.Run through each of the six factors and give GPT Image 2 something concrete for it:
- Subject. Not "a woman" but "a woman in her 60s with silver hair and laugh lines." Specificity is what separates a real person from a stock composite.
- Composition & framing. Name the shot size, angle, and placement — "tight head-and-shoulders, slightly low angle, subject off-center to the left." This is where the model's reasoning does the most work.
- Action. A subject mid-action ("pouring coffee," "laughing mid-sentence") reads as candid; a subject just posing reads as generated.
- Setting. Ground the scene: "a sunlit Lisbon café with worn marble tables," not "a café."
- Lighting. The single biggest realism lever. Name a real setup and its direction — see the table below.
- Style. Tell it the photographic register: editorial fashion, street documentary, food photography, shot on Kodak Portra 400.
Then the constraints line. Name the lens — the lens controls the whole feel, so an 85mm f/1.8 gives a flattering portrait with soft background falloff, a 35mm keeps a scene grounded and documentary, a 100mm macro gets extreme close texture, and a 24mm deep focus holds a whole landscape sharp. Add the aspect ratio and resolution plainly: "Aspect ratio 3:2, 4K." Requesting 2K or 4K matters because GPT Image 2 generates natively at that resolution rather than upscaling blindly, and image edges are multiples of 16.
| Factor | Concrete words that work |
|---|---|
| Subject | "a weathered fisherman in his 70s"; "a matte-black ceramic mug"; "a red 1967 coupe" |
| Composition | tight head-and-shoulders; three-quarter angle; overhead flat-lay; low camera looking up; subject off-center |
| Lens | 85mm f/1.8 (portrait); 35mm (street/scene); 100mm macro (detail); 24mm deep focus (interior/landscape) |
| Lighting | three-point softbox setup; golden-hour backlight with long shadows; chiaroscuro, harsh high contrast; soft overcast diffusion; rim light |
| Setting | "a rain-slicked Tokyo alley at night"; "a bright Scandinavian kitchen"; "a foggy pine ridge at dawn" |
| Style / texture | editorial fashion; candid documentary; shot on Kodak Portra 400; visible skin pores; fabric weave; condensation on glass; dust motes |
Keep the GPT Image 2 prompt cheat sheet open while you build these — it lists the supported aspect ratios, lens language, and lighting terms in one place.
10 photorealistic example prompts
Ten complete prompts, one per genre. Each is a full paragraph you can paste as-is into ChatGPT, then swap the specifics for your own subject.
1. Portrait
A candid close-up portrait of a woman in her early 60s with silver hair and soft laugh lines, glancing just off-camera with a faint smile. Tight head-and-shoulders framing, subject slightly off-center to the left, so the dark backdrop breathes on the right. Warm three-point softbox lighting with a soft key from the left and a subtle rim light separating her from the background. Editorial portrait style, natural color. Visible skin texture and pores, fine flyaway hairs catching the light. Shot on 85mm f/1.8, shallow depth of field with gentle background bokeh. Aspect ratio 4:5, 4K.Why it works: The 85mm lens plus a named softbox setup and explicit pore-level texture cues are exactly the details that kill the plastic look.
2. Product
A matte-black ceramic coffee mug centered on a smooth concrete surface, three-quarter angle, a thin curl of steam rising from it. Clean three-point softbox lighting with a soft gradient falloff on the seamless background and a crisp specular highlight down one edge of the mug. Premium e-commerce product-photography style, true material color. Fine matte-ceramic surface texture, a faint reflection on the concrete below. Shot on 100mm macro, moderately shallow depth of field. Aspect ratio 4:3, 4K.Why it works: Controlled softbox lighting plus a single specular highlight is how catalog shots read as premium; see more in the GPT Image 2 product photography prompts.
3. Food
An overhead shot of a rustic sourdough loaf, torn open to show an open, airy crumb, on a floured dark walnut board. A linen napkin, scattered flour, and a small dish of butter sit beside it. Soft overcast window light from the left with gentle falloff, no harsh highlights. Natural food-photography style. Crisp, blistered crust texture, visible flour dust, a faint sheen on the butter. Shot on 50mm, slight top-down angle. Aspect ratio 1:1, 4K.Why it works: Soft overcast diffusion is the flattering light real food photographers use, and the crust-and-flour texture cues make it look edible rather than rendered.
4. Street
A candid street scene of a man in a wool overcoat crossing a rain-slicked Tokyo alley at night, mid-stride, umbrella tilted against the drizzle. Framed at a slight angle to capture the full alley, neon signs reflected in the wet asphalt. Low-key lighting from mixed neon and a single overhead lamp, high contrast with deep shadows. Gritty street-documentary style, shot on film. Rain streaks visible in the light, texture in the wet pavement and the coat's weave. Shot on 35mm. Aspect ratio 3:2, 4K.Why it works: The 35mm framing and mixed practical lighting give it the grounded, unstaged feel real street photography has.
5. Interior
A bright Scandinavian living room with a linen sofa, pale oak floors, and a single large potted olive tree by tall windows. Wide interior view, straight-on and level to keep the vertical lines true. Soft morning daylight pouring through sheer curtains, gentle diffusion and long soft shadows across the floor. Architectural-digest interior style. Visible fabric weave on the sofa, grain in the oak, dust motes drifting in the light beam. Shot on 24mm, deep focus. Aspect ratio 16:9, 4K.Why it works: A wide level lens keeps walls from bowing, and diffused daylight with dust in the beam is the signature of real interior photography.
6. Landscape
A foggy pine ridge at dawn, layered mountains fading into pale mist behind it, a single hawk soaring in the distance. Wide landscape framing, deep depth of field so the whole scene is sharp front to back. Golden-hour backlight breaking over the ridge with long shadows and warm rim light on the treetops, cool blue shadow in the valley. Fine-art landscape style. Crisp needle detail on the nearest pines, soft atmospheric haze in the distance. Shot on 24mm deep focus. Aspect ratio 16:9, 4K.Why it works: Golden-hour backlight plus atmospheric haze gives the depth and warmth that flat, evenly-lit AI landscapes miss.
7. Macro
An extreme close-up of a single water droplet clinging to the edge of a green leaf, the veins of the leaf refracted inside the drop. Framed so the droplet fills the center, background dissolving into smooth green bokeh. Soft diffused side lighting picking out the droplet's rim and the tiny hairs on the leaf surface. Natural macro-photography style. Razor-sharp focus on the droplet, visible surface tension, dew and fine leaf texture. Shot on 100mm macro, extremely shallow depth of field. Aspect ratio 3:2, 4K.Why it works: A 100mm macro with a paper-thin plane of focus is exactly how real macro looks — one sharp point and everything else melting away.
8. Editorial fashion
A high-fashion editorial shot of a model in a structured crimson wool coat, standing against a raw concrete wall, chin lifted, one hand in her pocket. Full-length framing, subject centered with headroom, low camera looking slightly up. Dramatic chiaroscuro lighting, a single hard key from the side carving deep shadows and harsh high contrast. Vogue-style editorial fashion, shot on medium-format film. Rich fabric weave and structure in the coat, sharp skin and fabric texture, deliberate grain. Shot on 85mm. Aspect ratio 4:5, 4K.Why it works: Chiaroscuro's hard single key and deep shadow is a real studio technique, and it reads as intentional editorial rather than flat and generic.
9. Automotive
A glossy red 1967 coupe parked on a coastal road at dusk, three-quarter front angle, the ocean and a fading orange sky behind it. Low camera position looking slightly up at the car, the road leading out of frame. Golden-hour backlight rimming the roofline with warm long reflections streaking down the paint, cool ambient fill on the shadow side. Cinematic car-commercial style. Deep reflective paint with visible highlights, chrome catching the sky, fine road texture. Shot on 50mm. Aspect ratio 16:9, 4K.Why it works: Cars live or die on reflections; golden-hour rim light streaking down the paint is what makes the surface read as real glossy metal.
10. Candid documentary
A candid documentary frame of a grandfather and his young granddaughter baking together in a warm kitchen, both mid-laugh, flour on their hands and the countertop. Medium shot, natural eye-level angle, the two framed close with the kitchen falling off behind them. Soft window daylight from the right with a warm ambient bounce, gentle contrast. Photojournalistic documentary style, shot on film. Real skin texture, flour dust hanging in the light, worn wooden countertop and fabric detail. Shot on 35mm, shallow depth of field. Aspect ratio 3:2, 4K.Why it works: Two subjects caught mid-action in soft natural light is the anatomy of an honest candid photo — nobody is posing at the camera, which is exactly the relationship Native Thinking Mode reasons out.
Mistakes that break realism
Most fake-looking GPT Image 2 output traces back to one of these habits. Fix them before you touch anything else.
- Tag-soup instead of sentences. "woman, cafe, 85mm, bokeh, golden hour, moody" throws away every relationship in the scene — and Native Thinking Mode has nothing to reason about. Write it as a paragraph a photographer could read.
- Over-stacking adjectives. Ten adjectives on one noun ("a stunning gorgeous beautiful ethereal dramatic cinematic portrait") cancel each other out. Pick a few specific, concrete ones.
- No lighting at all. If you don't name the light, you get flat, sourceless illumination — the single most common tell of an AI image. Always name a setup and a direction.
- Quality-word spam. "8k ultra realistic hyperdetailed masterpiece award-winning" pushes toward the generic AI aesthetic, not away from it. Cut it; ask for "2K" or "4K" once and let the real detail come from your description.
- Ignoring aspect ratio. Leaving it out gives you a default crop that rarely fits. State "Aspect ratio X:Y" every time; GPT Image 2 handles 3:1 down to 1:3.
- No reference for consistency. Expecting the same face or product across shots from text alone won't work — attach a reference image and reuse it.
- Editing from scratch. When an image is nearly right, regenerating throws away what worked. Edit conversationally instead — see below.
Refine with conversational edits
When an image lands around 80% right, don't reroll — attach it, describe only the change, and lock everything else. GPT Image 2 follows edit instructions tightly and keeps faces and products consistent, so it holds the composition and lighting you already liked and adjusts just the one thing you named, at the original aspect ratio.
Keep the pose, framing, lighting, and background exactly the same. Change only the model's coat from crimson to deep forest green, same wool texture and structure. Preserve the original aspect ratio and resolution.Keep the subject, lens, and composition identical. Change the lighting from soft overcast to warm golden-hour backlight coming from behind the loaf, with longer shadows across the board. Do not alter anything else.Naming what stays identical is the whole trick — it stops the model from redrawing the face or the frame while it makes your one edit. For more paste-ready starting points, browse the full roundup, keep the cheat sheet handy, or, if you also use Google's model, compare notes in the guide to prompting Nano Banana for photorealism.
Frequently Asked Questions
Why does GPT Image 2 reward specific, complete prompts?
GPT Image 2 uses Native Thinking Mode — it reasons about composition, object counts, lighting, and constraints before it renders the first pixel. That reasoning step means a specific, complete natural-language brief needs far fewer rerolls, because the model works out the relationships in your scene up front. Vague or tag-style prompts give it less to reason about, so it lands on a generic average. Describe the shot the way you would brief a photographer, in full sentences, and let the model plan the frame.
What aspect ratio and resolution should I use for realistic images?
Match the aspect ratio to the shot: 3:4 or 4:3 for portraits and product, 16:9 for landscapes, 9:16 for phone verticals, 1:1 for social, and up to 3:1 or 1:3 for ultra-wide or tall crops. State it plainly as "Aspect ratio 16:9". GPT Image 2 renders natively up to 2K (about 2560×1440), roughly double GPT Image 1.5, with optional 4K upscaling. Ask for "2K" for screen use and "4K" when you need print detail. Image edges are multiples of 16.
How do I avoid the plastic, over-smoothed AI look?
The plastic look comes from missing lighting and missing texture. Name a real lighting setup (three-point softbox, golden-hour backlight, chiaroscuro, soft overcast) and add explicit texture cues — visible skin pores, fine flyaway hairs, fabric weave, condensation on glass, dust in the light. Drop spammy quality words like "8k ultra realistic hyperdetailed"; they push toward the generic AI aesthetic. Full descriptive sentences with one clear light source almost always beat a pile of adjectives.
How do I add believable text to a photorealistic image?
Text rendering is GPT Image 2's headline strength — roughly 99% accuracy in English and strong results in Chinese, Japanese, Korean, Hindi, Bengali, and Arabic — so it usually renders clean. Still, help it: wrap the exact words in quotes, keep each string short (1 to 4 words in ALL-CAPS is most reliable), name the font and weight, and put the text instruction near the start of the prompt. Say where it sits and on what surface, such as the word 'FRESH' in bold condensed sans-serif embossed on a metal lid.
How do I keep a face or product consistent across several images?
Use reference images. GPT Image 2 accepts reference or input images and holds faces and products consistent across generations. Attach a clear photo, describe the new scene, and state explicitly what must stay identical — the identity, the label, the shape. For a series, reuse the same reference each time and change only the setting, wardrobe, or lighting in the prompt so the subject stays recognizable shot to shot.
Is GPT Image 2 output commercially usable, and is it watermarked?
GPT Image 2 images carry C2PA content-credential metadata marking them AI-generated, and there is no large visible watermark on standard outputs. Commercial use follows OpenAI's usage terms — you own the images you create subject to those terms, so review the terms tied to your account before shipping paid or branded work. As with any generative model, review every image for accuracy, likeness, and brand-safety before you publish.
When should I edit conversationally instead of regenerating?
Once an image is roughly 80% right, edit it rather than rolling the dice again. GPT Image 2 supports conversational editing: attach the image, describe only the change, and list what must stay identical — the pose, lighting, and background exactly the same, change only the jacket to deep red. Regenerating from scratch throws away the composition and lighting you already liked and gives you a different face and frame. Conversational edits preserve what worked and adjust just the one thing you named, at the original aspect ratio.