Realistic Wan 2.7 video comes from writing like a director's shot brief, not a wish list. Name one subject, one scene, one primary motion, specific light, a single camera move, a style, and matched native audio, and the clip reads as filmed rather than generated. Wan (Tongyi Wanxiang) is Alibaba's open-weight model, shipped under an Apache 2.0 license so the output is free to use commercially, and it does text-to-video, image-to-video, and synced sound. This guide gives you the exact formula, then 10 finished prompts you can paste into Alibaba Cloud Model Studio, a hosted API, or a local run. For a broader library once the pattern clicks, see the 35 best Wan prompts.

Advertisement

The 6-part Wan formula

Every strong Wan shot is one paragraph built from six beats, in this order, with the audio lines underneath. Write each beat once and be concrete — Wan rewards specificity and punishes vagueness. Keep the whole thing under about 200 words so the model does not average competing ideas together.

BeatWhat to write
SubjectWho or what is on screen — appearance, wardrobe, age, material. "A weathered fisherman in a yellow oilskin," not "a man."
ScenePlace plus time of day, so the light and mood have a reason to exist.
MotionOne clear primary motion the shot is about — not three stacked verbs.
LightingA named lighting term — golden hour, blue hour, low-key, softbox, volumetric light.
CameraShot size plus exactly one move — slow dolly-in, static locked-off, handheld follow, orbit clockwise.
StyleLens and finish — 35mm, shallow depth of field, film grain, or a look reference.

Assembled, a shot reads like this: "A weathered fisherman in a yellow oilskin hauls a rope over the gunwale of a small boat. Open sea at dawn, grey swell. Slow dolly-in following the pull. Overcast soft light, volumetric mist. Medium shot, 35mm, film grain. SFX: rope creaking, water slapping the hull. Ambient: gulls, distant wind." That is the whole method — the rest of this guide sharpens each beat. Grab ready vocabulary from the Wan prompt cheat sheet when you need shot sizes and moves fast, or start from a filled skeleton with the Wan prompt templates.

One primary motion — the #1 rule

Wan follows a motion hierarchy. A single clear primary motion renders clean on most generations; stack several equal motions and the shot turns to chaos as the model fights to serve them all at once. So each shot gets one thing that moves the most — a walking subject, falling snow, pouring water, a rotating platform, an opening lid — and everything else stays subordinate to it.

This is the single highest-leverage fix for the fake look. If your idea has a car drifting and a crowd cheering and a drone swooping, pick the one that carries the shot and demote the rest to ambient detail. Need a genuine second beat? Make it a second clip, or use Wan's multi-shot mode — label each scene with its own camera grammar and a sound line, and let the model cut between them instead of blending them into one frame.

Camera & lens language Wan understands

The AI look usually comes from motion no real operator would produce and light with no source. Fix both with plain film vocabulary Wan was trained to recognize.

Camera moves: pick one. Wan responds well to slow dolly-in and push-in for emphasis, a static / locked-off wide shot that lets the action breathe, a handheld follow for documentary energy, crane up for a reveal, orbit clockwise around a product, rack focus from foreground to background, and slow pan or tilt down to travel across a scene. Avoid drifting, floating moves with no anchor — they are the classic tell.

Lens & lighting terms: name a focal length and a light quality. Lenses run 24mm and 35mm for wide documentary work, 50mm and 85mm for portraits, 100mm for macro. Lighting terms Wan reads cleanly include golden hour, blue hour, low-key, high-key, softbox, and volumetric light. Add shallow depth of field and film grain for texture. "Cinematic lighting" alone is too vague — Wan needs a direction and a quality to aim at.

Advertisement

Writing native audio: voiceover, SFX, music

Native synced audio is one of Wan's headline features: it generates dialogue, sound effects, ambient sound, and music together with the picture. Silent or generically-scored clips read as AI instantly, so describe sound on every prompt using labeled lines under the shot:

  • Dialogue / Voiceover. Put speech on its own line and write it in the language you want spoken — Wan is strong in Mandarin and English. Keep it to one or two short lines for a clip this length: Dialogue: the barista says, "Morning! The usual oat flat white?"
  • SFX, spelled out. Name the specific sounds the motion makes: SFX: espresso machine hissing, cups clinking.
  • Ambient bed. Set the background layer: Ambient: low café chatter, soft jazz.
  • Music. Describe the score by mood and instrument: Music: sparse, elegant piano.

Match the audio to the shot size. A close-up implies intimate, present sound; a wide establishing shot implies distant, reverberant ambience. When speech and SFX line up with what the camera shows, the brain accepts the clip as real. In image-to-video you can go further and pass a real audio file, so Wan syncs lip and body motion to an actual voice track instead of generating one.

Image-to-video & start/end frames

Wan's image-to-video is one of its best tricks: upload a start frame and describe only the motion, camera behavior, and mood you want added to it. The rule that keeps it realistic is to anchor the source — tell Wan to keep the subject, colors, logo, and style identical to the reference image, and keep the described change small and physical (a blink, a breeze, a slow rotation). Skip that anchor and the subject drifts and morphs off-model.

Wan 2.7 also accepts an optional end frame, so you can define both the first and last image and let the model generate the motion between them — ideal for matching an edit or hitting a precise product hero pose. Pair this with the audio-file input above and you can animate a still portrait that lip-syncs to a real voiceover.

Negative prompts & settings

Two categories of control live outside the prompt paragraph, in the app or API — not inline in the text:

  • Negative prompt. A separate field where you list what to suppress — for example blurry, warped hands, extra fingers, text, watermark, flicker, oversaturated. Put unwanted things here, not as "no X" inside the main prompt.
  • Settings. Resolution (up to 1080p), aspect ratio (16:9, 9:16, 1:1), duration, seed, and prompt expansion / "thinking mode" are all chosen in the interface. Wan has no --ar-style parameters. A trailing note like (16:9, 1080p, 10s) is just a reminder of what to select — never paste it into the prompt box expecting it to do anything.

7 mistakes that make Wan look fake

Most unrealistic Wan clips fail on the same handful of errors. Check every prompt against this list before you generate.

  • Stacking multiple motions. Several equal motions in one shot produce chaotic, unstable frames. Choose one primary motion and demote the rest.
  • Vague subjects. "A person" or "a car" gives Wan nothing to hold. Name appearance, wardrobe, and material so the subject stays consistent.
  • Conflicting styles. "Photorealistic anime" or "oil-painted 4K photo" pulls the model two ways. Pick one visual language and commit to it.
  • No lighting term. Missing light is a top tell. Always name a source and quality — golden hour, low-key, softbox, volumetric light.
  • Overlong prompts. Past ~200 words Wan averages competing ideas and the shot goes muddy. Trim to the six beats plus audio.
  • Forgetting to anchor the source in image-to-video. Without "keep the subject and colors identical," the reference drifts and morphs. Lock it every time.
  • Writing settings inside the prompt. Aspect ratio, duration, and negatives are app or API controls. Putting "16:9, no blur" in the text does nothing useful — use the fields.

Fix these seven and your hit rate on realistic takes climbs sharply. When you are ready for a bigger set of finished shots, work through the 35 best Wan prompts.

10 example prompts

Each prompt below is a complete six-beat shot with a described audio bed, a Why it works line, and a plain settings note. Resolution, aspect ratio, and duration are settings — the note tells you what to pick, not text to paste into the prompt.

1. Café barista (dialogue)

A young barista in a green apron with rolled-up sleeves steams milk behind a busy espresso bar and glances up with a quick smile. Small specialty café, mid-morning light through a front window. Slow dolly-in on the barista. Warm practical lighting, shallow depth of field. Medium shot, 35mm.
Dialogue: the barista says, "Morning! The usual oat flat white?"
SFX: espresso machine hissing, cups clinking.
Ambient: low café chatter, soft jazz.
(16:9, 1080p, 8s)

Why it works: one short line, one gentle move, and synced SFX that match exactly what the hands are doing.

2. Rain-slicked city walk (mood)

A woman in a long charcoal coat walks briskly down a wet city sidewalk at night, collar up against the drizzle. Downtown street, neon signage reflected in the puddles. Handheld follow tracking alongside her at a steady pace. Low-key lighting, practical neon, volumetric haze. Wide shot, 35mm, film grain.
SFX: footsteps on wet pavement, a distant car horn.
Ambient: light rain, muffled traffic hum.
(9:16, 1080p, 8s)

Why it works: a single motivated tracking move and reflective practicals give real depth without any camera drift.

3. Product hero — wireless earbuds (ad)

A matte-black earbud case rests on a brushed-concrete slab, the lid slowly opening to reveal the earbuds and a soft glow inside. Minimal studio set. Static locked-off camera as the lid opens. Softbox top light with a subtle rim, shallow depth of field. Extreme close-up, 85mm.
SFX: a crisp magnetic click as the lid opens.
Ambient: near silence, faint studio room tone.
Music: minimal upbeat electronic.
(16:9, 1080p, 6s)

Why it works: a locked-off frame plus one precise SFX cue keeps focus on the product; start from a real product photo for exact branding.

4. Mountain establishing shot (nature)

A lone hiker crests a rocky ridgeline as morning fog drifts through the valley below. Alpine range at sunrise. Slow crane up rising to reveal the full valley. Golden-hour light, volumetric light through the fog. Extreme wide shot, 24mm.
SFX: gusting wind, gravel underfoot.
Ambient: distant birdsong, open-air stillness.
Music: sweeping cinematic strings.
(16:9, 1080p, 8s)

Why it works: the crane is the only move, so the reveal stays smooth and the scale reads as filmed.

5. Kitchen how-to (creator / Shorts)

A home cook cracks an egg one-handed into a steel bowl, then whisks briskly as flour dusts the counter. Bright home kitchen, midday. Static camera framed overhead on the bowl. High-key soft daylight, shallow depth of field. Close-up, 50mm.
Voiceover (calm narrator): "See how fast that comes together?"
SFX: shell cracking, whisk clinking against steel.
Ambient: quiet kitchen, a faint radio.
(9:16, 1080p, 8s)

Why it works: one clear action, tight framing, and diegetic sound make a vertical clip feel like a real tutorial.

6. Vintage motorcycle reveal (action)

A vintage motorcycle rolls out of a dark tunnel onto an open desert highway at dusk, the rider silhouetted against the light. Empty road, blue-hour sky fading to orange. Slow push-in as the bike approaches. Backlight from the low sun, film grain. Full shot, 35mm.
SFX: engine growl building, tyres over grit.
Ambient: dry desert wind.
Music: driving low percussion.
(16:9, 1080p, 8s)

Why it works: backlight and a single push-in sell the reveal while the audio ramp matches the acceleration.

7. Workshop interview (talking head)

A middle-aged woodworker in a canvas apron speaks directly to camera, sawdust on her forearms, workshop tools blurred behind her. Small workshop, warm afternoon light. Static camera with a very slow push-in. Soft window key with a practical work-lamp fill, shallow depth of field. Close-up, 50mm.
Dialogue: she says, "I've been building tables for twenty years."
SFX: a faint saw hum in the background.
Ambient: quiet workshop room tone.
(16:9, 1080p, 6s)

Why it works: a nearly still frame and one short line keep Wan's synced lip movement clean and the delivery natural.

8. Skincare serum pour (beauty ad)

A glass dropper releases a single droplet of clear serum onto the back of a hand, the drop spreading slowly across the skin. Clean studio surface, soft pastel backdrop. Slow push-in on the droplet. Diffused softbox light with a subtle rim, shallow depth of field. Extreme close-up, 100mm macro.
SFX: a soft liquid tap as the drop lands.
Ambient: calm studio silence.
Music: airy ambient pad.
(9:16, 1080p, 6s)

Why it works: a macro push-in and a single delicate SFX give a premium, tactile feel that reads as real product footage.

9. Coastal B-roll (cinematic)

Waves roll over black volcanic rocks as sea spray catches the low sun. Rugged coastline, late golden hour. Slow static hold with a faint handheld drift. Golden-hour backlight, anamorphic-style flares, film grain. Wide shot, 35mm.
SFX: waves crashing, spray hissing over rock.
Ambient: steady ocean roar, distant gulls.
(16:9, 1080p, 8s)

Why it works: falling water is a continuous motion Wan tracks cleanly, and a subtle drift adds life without breaking realism.

10. Portrait photo comes alive (image-to-video)

[Start frame: a studio portrait of a young woman]
Animate the portrait: she blinks softly, a gentle smile forms, and a light breeze lifts a few strands of her hair. Keep her features, wardrobe, and colors identical to the source image. Static camera with a very slow push-in. Preserve the original lighting and grade.
SFX: none.
Ambient: quiet room tone.
(4:5, 1080p, 5s)

Why it works: anchoring the source and describing only subtle motion — a blink, a breath, a breeze — keeps the person on-model instead of morphing.

Frequently Asked Questions

What makes Wan 2.7 video look realistic?

Three things: one clear primary motion per shot, a single motivated camera move written in real film language, and native audio described to match the scene. Vague subjects, missing lighting terms, and several equal motions stacked together are what usually produce the fake AI look.

What is the Wan prompt formula?

Write the shot in this order: Subject + Scene + Motion + Lighting + Camera + Style, then add audio lines (Dialogue, Voiceover, SFX, Ambient, Music). Keep the whole prompt under about 200 words so Wan does not average competing ideas together.

How do I write native audio for Wan?

Add labeled lines under the shot: Dialogue: or Voiceover: for speech (in the language you want spoken), SFX: for specific sound effects, Ambient: for the background bed, and Music: for score. In image-to-video you can also pass a real audio file so Wan syncs lip and body motion to it.

How many motions should one Wan shot have?

One primary motion. Wan follows a motion hierarchy — a single clear motion renders clean, while stacking several equal motions turns the shot to chaos. If you need more, split it into a second clip or use Wan's multi-shot mode with labeled scenes.

Can I set aspect ratio or duration inside a Wan prompt?

No. Wan has no inline parameters like Midjourney's --ar. Resolution (up to 1080p), aspect ratio (16:9, 9:16, 1:1), duration, seed, and negative prompt are all chosen in the app or API. A trailing note like (16:9, 1080p, 10s) is just a reminder of what to select, not text to paste.

How does Wan image-to-video work?

Upload a start frame and describe only the motion, camera behavior, and mood to add to it. Anchor the source by telling Wan to keep the subject, colors, and style identical. Wan 2.7 also accepts an optional end frame, so you can define both the first and last image and let the model fill the motion between them.

Do I still describe audio if Wan generates it automatically?

Yes. Wan generates synced dialogue, SFX, and ambient sound, but it follows your description far more reliably than its guesses. Add an SFX line and an Ambient line to every prompt — described audio is one of the biggest drivers of perceived realism.

Is Wan free to use commercially?

Yes. Wan 2.7's weights ship under an Apache 2.0 license, which permits commercial use without licensing fees. If you run it through a hosted service such as Alibaba Cloud Model Studio or a third-party API, that platform's pricing applies, but the output is yours to use.

Advertisement