Keep this page open while you write Wan 2.7 prompts. It has the exact vocabulary the model understands — shot sizes, camera moves, lighting terms, and the native-audio syntax that makes Wan generate synced dialogue and sound. Wan (Tongyi Wanxiang) is Alibaba's open-weight video model, and because it ships under an Apache 2.0 license the clips are free to use commercially. Every table below is copy-paste ready, and the settings section tells you what belongs in the app rather than the prompt. For full walkthroughs, see the 35 best Wan prompts and how to prompt Wan for realistic video.

Advertisement

The 6-part shot formula

Wan responds best when you write like a director handing a crew a shot brief — one flowing paragraph, then the audio lines beneath it. Every strong shot breaks into six beats, in this order. The last one, style, is where the audio lines live, and native audio is what sets Wan 2.7 apart from older open models.

BeatWhat to write
SubjectWho or what is on screen — specific appearance, wardrobe, materials.
ScenePlace plus time of day.
MotionOne clear primary motion. Not three actions stitched together.
LightingOne lighting term that sets the mood.
CameraShot size plus one movement.
StyleLens or film reference, then the audio lines beneath — Dialogue:, SFX:, Ambient:.

Stack the first five beats into one paragraph, add the audio lines underneath, and set duration, resolution, and aspect ratio in the app before you generate. Grab terms from the tables below and drop them straight in. Keep the whole prompt under about 200 words so the model does not average competing ideas. See Wan prompt templates for fill-in-the-blank skeletons built on this exact formula.

Shot sizes

Shot size controls how much of the subject and setting is visible. Pick one per shot and name it explicitly.

TermWhat it framesUse for
Extreme wide shotSubject tiny against a huge environmentEstablishing scale, landscapes, openers
Wide shotFull subject plus surrounding spaceShowing action and location together
Full shotSubject head to toe, little backgroundFull-body movement, dance, sport
Medium shotWaist upDialogue, everyday action, product use
Medium close-upChest and headTalking-head clips, reactions
Close-upHead and shouldersEmotion, spoken lines
Extreme close-up / macroEyes, hands, or a small detail fills the frameTexture, product detail, water droplets
Over-the-shoulderForeground shoulder/head, subject beyondConversations, listener's POV
POVCamera is the subject's eyesImmersive first-person action
Top-down / overheadLooking straight down on the sceneFood, unboxing, flat-lay reveals

Camera moves — one per shot

Use exactly one move per shot. Wan follows a motion hierarchy, so stacking moves — a pan that also dollies and orbits — makes the motion unstable and shaky. Need a second move? Make it a second clip or use Wan's multi-shot mode with labeled scenes.

MoveEffectWhen to use
Static / locked-offNo camera motion at allDialogue, calm product shots, stillness
PanPivots horizontally from a fixed pointRevealing a wide space, following lateral movement
TiltPivots vertically from a fixed pointRevealing height — a building, a tall subject
Dolly in / outPhysically moves toward or away from subjectBuilding intensity (in) or releasing it (out)
TrackingMoves alongside a moving subjectFollowing a walk, run, or vehicle at constant distance
Crane / jibRises or descends smoothly, often changing angleGrand reveals, rising above a scene
Orbit / arc (clockwise)Circles around the subjectHero shots, product turnarounds
Push-inSlow, subtle move closer without a full dollyQuiet emotional emphasis
Pull-backSlow move away to reveal contextEnding a scene, revealing scale
Handheld followSlight natural shake, human-operated feelDocumentary realism, urgency, candid clips
Rack focusShifts focus from one plane to anotherDirecting the eye, reveals within a shot
Whip-panVery fast horizontal pan, motion-blurredEnergetic transitions, action beats

Lighting & lens

Pair one lighting term with one lens or style term for a consistent, filmic look. Two is plenty — more and Wan starts averaging them out.

TermLook
Golden hourWarm, low-angle sun, long soft shadows
Blue hourCool twilight tones just after sunset
High-keyBright, even, low-contrast — commercial, upbeat
Low-keyDark, high-contrast, moody shadows
SoftboxDiffused, flattering studio light
Practical lightsIn-scene sources — lamps, neon, candles
Volumetric lightVisible beams and haze — fog, dust, god rays
Backlight / rim lightLight behind subject, glowing edge outline
Shallow depth of fieldSharp subject, soft blurred background
24mmExpansive wide-angle, slightly exaggerated perspective
35mmNatural, documentary feel
50mmNeutral, close to how the eye sees
85mmCompressed, flattering portrait feel
100mm macroExtreme magnification of a small subject
Anamorphic flaresWidescreen look with horizontal lens flares
Film grainSubtle texture, analog feel
Advertisement

Native audio syntax

This is Wan 2.7's headline feature: it generates dialogue, voiceover, sound effects, ambient sound, and music synced to the video. Write each layer as its own line beneath the visual paragraph. Keep dialogue short — a clip is only 5-15 seconds, so aim for one or two brief lines.

LayerSyntaxWhat to write
DialogueDialogue: prefixDialogue: the barista says, "One oat flat white, coming up." — attribute the speaker, keep it to a line or two.
VoiceoverVoiceover: prefixVoiceover (calm female narrator): "Start with ripe tomatoes." — off-screen narration.
SFXSFX: prefixSFX: espresso machine hissing, cups clinking. — name each discrete sound.
AmbientAmbient: prefixAmbient: quiet café murmur, soft street noise. — the background bed of the scene.
MusicMusic: prefixMusic: sparse, elegant piano. — set mood and genre, not a specific track.

In image-to-video you can also pass a real audio file alongside the start frame, and Wan matches lip and body motion to the track instead of generating the sound itself.

Output settings & modes

These are chosen in the app or API (Alibaba Cloud Model Studio or a hosted service) before you generate — not written into the prompt text. Wan has no inline parameters like Midjourney's --ar, so typing a duration or aspect ratio in the prompt does nothing.

SettingOptionsNote
ResolutionUp to 1080p at 24fpsDraft low, render the final take at 1080p.
Aspect ratio16:9, 9:16, 1:1, 4:516:9 for widescreen; 9:16 for Shorts, Reels, TikTok; 1:1 and 4:5 for feed.
Duration~5-15s per generationKeep the primary motion simple so it reads cleanly across the whole clip.
SeedAny integerReuse a seed to reproduce or iterate on a result.
Negative promptFree textList what to avoid — extra limbs, text artifacts, warping.
ModesText-to-video, image-to-video, multi-shotImage-to-video takes a start frame and an optional end frame; multi-shot chains labeled scenes.
Prompt expansion / thinking modeOn / offLets Wan enrich a short prompt. Leave off when you want tight control over a specific shot.

Example prompts

Five complete Wan shots built from the tables above. Each ends with a plain settings note in parentheses — a reminder of what to pick in the app, not text to paste into the prompt.

1. Product close-up with dialogue

A barista in a green apron sets a ceramic cup on a wooden counter and pushes it toward the camera, steam curling upward. Cozy café interior, early morning light through a tall window. Medium close-up, slow push-in. Golden hour lighting, shallow depth of field, 85mm.
Dialogue: the barista says, "One oat flat white, coming up."
SFX: espresso machine hissing, cup settling on wood.
Ambient: quiet café murmur, soft jazz.
(16:9, 1080p, 8s)

Best for: product and café ads where the spoken line sells it. Why it works: one action plus one move keeps the motion stable while Wan's synced audio does the selling.

2. Walking city scene

A woman in a red raincoat walks briskly down a rain-slicked sidewalk at night, weaving between umbrellas. Downtown street, neon signs reflecting in puddles. Wide shot, handheld follow alongside her. Low-key lighting, practical neon, 35mm, film grain.
SFX: footsteps splashing on wet pavement, a distant car horn.
Ambient: steady rain, muffled city traffic.
(9:16, 1080p, 10s)

Best for: vertical social openers and moody b-roll. Why it works: one motion, one move, and an ambient bed that grounds the whole shot.

3. Nature establishing shot

A lone hiker crosses a ridgeline as morning fog drifts through the valley below. Mountain range at sunrise. Extreme wide shot, slow crane rising to reveal the full valley. Golden hour lighting, volumetric light through the fog, 24mm.
SFX: wind gusting across the ridge, a single crow call.
Ambient: distant birdsong, a faint river below.
Music: sweeping cinematic strings.
(16:9, 1080p, 10s)

Best for: openers and travel content. Why it works: a single steady crane keeps an epic reveal coherent where a swooping move would wobble.

4. Two-person dialogue

Two friends sit across a small table, one leaning in as the other reacts. Warm bistro interior, afternoon light through a large window. Over-the-shoulder shot, static locked-off camera. High-key lighting, softbox fill, shallow depth of field, 50mm.
Dialogue: she says, "You actually did it — you quit today?" He grins, "This morning."
SFX: cutlery clinking, a chair scraping nearby.
Ambient: low bistro chatter, warm background music.
(16:9, 1080p, 8s)

Best for: character beats and skits. Why it works: two short lines fit the clip, and a locked-off camera lets Wan spend its motion budget on synced lips.

5. Image-to-video product turntable

[Start frame: a clean product photo of a running shoe on a white background]
Add motion: the shoe rotates slowly clockwise on an invisible turntable while the studio reflection shifts across its surface. Keep the shoe, colors, and logo exactly as in the source image. Medium close-up, static camera. Preserve the clean white background and soft shadow.
SFX: a subtle low hum.
Music: minimal upbeat electronic.
(1:1, 1080p, 6s)

Best for: turning one catalog photo into a rotating clip with accurate branding. Why it works: anchoring "keep the logo exactly as in the source" stops Wan from redrawing the product while it adds motion.

Want more finished examples before you start? Browse the 35 best Wan prompts, read how to prompt Wan for realistic video, or grab ready-made skeletons from Wan prompt templates.

Frequently Asked Questions

What is the Wan prompt formula?

Write six beats in order: Subject, Scene (place plus time), Motion (one primary motion), Lighting, Camera (shot size plus one move), and Style — then add audio lines beneath. Stack the first beats into one flowing paragraph, like a shot brief handed to a crew, and the audio lines carry Wan's native sound.

Does Wan generate sound?

Yes. Wan 2.5 added native audio and Wan 2.7 keeps it: write Dialogue, Voiceover, SFX, Ambient, and Music lines beneath the visual paragraph and the model generates them synced to the picture. In image-to-video you can also pass a real audio file so Wan matches lip and body motion to the track.

How many camera moves should one Wan shot have?

One. Wan follows a motion hierarchy, so a single clear primary motion plus one camera move gives clean results. Stacking several equal moves — a pan that also dollies and orbits — makes the shot chaotic. Need a second move? Make it a second clip or use Wan's multi-shot mode with labeled scenes.

What resolution and length does Wan support?

Wan 2.7 outputs up to 1080p at 24fps, with clips generally in the 5-to-15-second range depending on the interface. Resolution, aspect ratio, duration, seed, and negative prompt are chosen as settings in the app or API — they are not written inline in the prompt text.

Can I set aspect ratio or length inside the prompt?

No. Wan has no inline parameters like Midjourney's --ar. The settings note in parentheses at the end of each prompt — such as (16:9, 1080p, 10s) — is a reminder of what to pick in the app or API, not text to paste into the prompt box.

How does Wan image-to-video work?

Upload a start frame and describe only the motion, camera behavior, and mood you want added to it. Wan 2.7 also accepts an optional end frame, so you can define both the first and last image and let the model generate the motion between them. Keep the described change small and physical to stay on-model.

Is Wan free to use commercially?

The Wan model weights are released under Apache 2.0, which permits commercial use without licensing fees. If you run it through a hosted service such as Alibaba Cloud Model Studio or a third-party API, that platform's pricing and terms still apply, but the output itself is yours to use commercially.

What is prompt expansion or thinking mode in Wan?

Prompt expansion (sometimes called thinking mode) lets Wan rewrite a short prompt into a richer, more detailed brief before generating. It is useful for filling in lighting and camera detail you left out, but for tight control over a specific shot, write the full six-beat prompt yourself and leave expansion off.

Advertisement