Keep this page open while you write Wan 2.7 prompts. It has the exact vocabulary the model understands — shot sizes, camera moves, lighting terms, and the native-audio syntax that makes Wan generate synced dialogue and sound. Wan (Tongyi Wanxiang) is Alibaba's open-weight video model, and because it ships under an Apache 2.0 license the clips are free to use commercially. Every table below is copy-paste ready, and the settings section tells you what belongs in the app rather than the prompt. For full walkthroughs, see the 35 best Wan prompts and how to prompt Wan for realistic video.
The 6-part shot formula
Wan responds best when you write like a director handing a crew a shot brief — one flowing paragraph, then the audio lines beneath it. Every strong shot breaks into six beats, in this order. The last one, style, is where the audio lines live, and native audio is what sets Wan 2.7 apart from older open models.
| Beat | What to write |
|---|---|
| Subject | Who or what is on screen — specific appearance, wardrobe, materials. |
| Scene | Place plus time of day. |
| Motion | One clear primary motion. Not three actions stitched together. |
| Lighting | One lighting term that sets the mood. |
| Camera | Shot size plus one movement. |
| Style | Lens or film reference, then the audio lines beneath — Dialogue:, SFX:, Ambient:. |
Stack the first five beats into one paragraph, add the audio lines underneath, and set duration, resolution, and aspect ratio in the app before you generate. Grab terms from the tables below and drop them straight in. Keep the whole prompt under about 200 words so the model does not average competing ideas. See Wan prompt templates for fill-in-the-blank skeletons built on this exact formula.
Shot sizes
Shot size controls how much of the subject and setting is visible. Pick one per shot and name it explicitly.
| Term | What it frames | Use for |
|---|---|---|
| Extreme wide shot | Subject tiny against a huge environment | Establishing scale, landscapes, openers |
| Wide shot | Full subject plus surrounding space | Showing action and location together |
| Full shot | Subject head to toe, little background | Full-body movement, dance, sport |
| Medium shot | Waist up | Dialogue, everyday action, product use |
| Medium close-up | Chest and head | Talking-head clips, reactions |
| Close-up | Head and shoulders | Emotion, spoken lines |
| Extreme close-up / macro | Eyes, hands, or a small detail fills the frame | Texture, product detail, water droplets |
| Over-the-shoulder | Foreground shoulder/head, subject beyond | Conversations, listener's POV |
| POV | Camera is the subject's eyes | Immersive first-person action |
| Top-down / overhead | Looking straight down on the scene | Food, unboxing, flat-lay reveals |
Camera moves — one per shot
Use exactly one move per shot. Wan follows a motion hierarchy, so stacking moves — a pan that also dollies and orbits — makes the motion unstable and shaky. Need a second move? Make it a second clip or use Wan's multi-shot mode with labeled scenes.
| Move | Effect | When to use |
|---|---|---|
| Static / locked-off | No camera motion at all | Dialogue, calm product shots, stillness |
| Pan | Pivots horizontally from a fixed point | Revealing a wide space, following lateral movement |
| Tilt | Pivots vertically from a fixed point | Revealing height — a building, a tall subject |
| Dolly in / out | Physically moves toward or away from subject | Building intensity (in) or releasing it (out) |
| Tracking | Moves alongside a moving subject | Following a walk, run, or vehicle at constant distance |
| Crane / jib | Rises or descends smoothly, often changing angle | Grand reveals, rising above a scene |
| Orbit / arc (clockwise) | Circles around the subject | Hero shots, product turnarounds |
| Push-in | Slow, subtle move closer without a full dolly | Quiet emotional emphasis |
| Pull-back | Slow move away to reveal context | Ending a scene, revealing scale |
| Handheld follow | Slight natural shake, human-operated feel | Documentary realism, urgency, candid clips |
| Rack focus | Shifts focus from one plane to another | Directing the eye, reveals within a shot |
| Whip-pan | Very fast horizontal pan, motion-blurred | Energetic transitions, action beats |
Lighting & lens
Pair one lighting term with one lens or style term for a consistent, filmic look. Two is plenty — more and Wan starts averaging them out.
| Term | Look |
|---|---|
| Golden hour | Warm, low-angle sun, long soft shadows |
| Blue hour | Cool twilight tones just after sunset |
| High-key | Bright, even, low-contrast — commercial, upbeat |
| Low-key | Dark, high-contrast, moody shadows |
| Softbox | Diffused, flattering studio light |
| Practical lights | In-scene sources — lamps, neon, candles |
| Volumetric light | Visible beams and haze — fog, dust, god rays |
| Backlight / rim light | Light behind subject, glowing edge outline |
| Shallow depth of field | Sharp subject, soft blurred background |
| 24mm | Expansive wide-angle, slightly exaggerated perspective |
| 35mm | Natural, documentary feel |
| 50mm | Neutral, close to how the eye sees |
| 85mm | Compressed, flattering portrait feel |
| 100mm macro | Extreme magnification of a small subject |
| Anamorphic flares | Widescreen look with horizontal lens flares |
| Film grain | Subtle texture, analog feel |
Native audio syntax
This is Wan 2.7's headline feature: it generates dialogue, voiceover, sound effects, ambient sound, and music synced to the video. Write each layer as its own line beneath the visual paragraph. Keep dialogue short — a clip is only 5-15 seconds, so aim for one or two brief lines.
| Layer | Syntax | What to write |
|---|---|---|
| Dialogue | Dialogue: prefix | Dialogue: the barista says, "One oat flat white, coming up." — attribute the speaker, keep it to a line or two. |
| Voiceover | Voiceover: prefix | Voiceover (calm female narrator): "Start with ripe tomatoes." — off-screen narration. |
| SFX | SFX: prefix | SFX: espresso machine hissing, cups clinking. — name each discrete sound. |
| Ambient | Ambient: prefix | Ambient: quiet café murmur, soft street noise. — the background bed of the scene. |
| Music | Music: prefix | Music: sparse, elegant piano. — set mood and genre, not a specific track. |
In image-to-video you can also pass a real audio file alongside the start frame, and Wan matches lip and body motion to the track instead of generating the sound itself.
Output settings & modes
These are chosen in the app or API (Alibaba Cloud Model Studio or a hosted service) before you generate — not written into the prompt text. Wan has no inline parameters like Midjourney's --ar, so typing a duration or aspect ratio in the prompt does nothing.
| Setting | Options | Note |
|---|---|---|
| Resolution | Up to 1080p at 24fps | Draft low, render the final take at 1080p. |
| Aspect ratio | 16:9, 9:16, 1:1, 4:5 | 16:9 for widescreen; 9:16 for Shorts, Reels, TikTok; 1:1 and 4:5 for feed. |
| Duration | ~5-15s per generation | Keep the primary motion simple so it reads cleanly across the whole clip. |
| Seed | Any integer | Reuse a seed to reproduce or iterate on a result. |
| Negative prompt | Free text | List what to avoid — extra limbs, text artifacts, warping. |
| Modes | Text-to-video, image-to-video, multi-shot | Image-to-video takes a start frame and an optional end frame; multi-shot chains labeled scenes. |
| Prompt expansion / thinking mode | On / off | Lets Wan enrich a short prompt. Leave off when you want tight control over a specific shot. |
Example prompts
Five complete Wan shots built from the tables above. Each ends with a plain settings note in parentheses — a reminder of what to pick in the app, not text to paste into the prompt.
1. Product close-up with dialogue
A barista in a green apron sets a ceramic cup on a wooden counter and pushes it toward the camera, steam curling upward. Cozy café interior, early morning light through a tall window. Medium close-up, slow push-in. Golden hour lighting, shallow depth of field, 85mm.
Dialogue: the barista says, "One oat flat white, coming up."
SFX: espresso machine hissing, cup settling on wood.
Ambient: quiet café murmur, soft jazz.
(16:9, 1080p, 8s)Best for: product and café ads where the spoken line sells it. Why it works: one action plus one move keeps the motion stable while Wan's synced audio does the selling.
2. Walking city scene
A woman in a red raincoat walks briskly down a rain-slicked sidewalk at night, weaving between umbrellas. Downtown street, neon signs reflecting in puddles. Wide shot, handheld follow alongside her. Low-key lighting, practical neon, 35mm, film grain.
SFX: footsteps splashing on wet pavement, a distant car horn.
Ambient: steady rain, muffled city traffic.
(9:16, 1080p, 10s)Best for: vertical social openers and moody b-roll. Why it works: one motion, one move, and an ambient bed that grounds the whole shot.
3. Nature establishing shot
A lone hiker crosses a ridgeline as morning fog drifts through the valley below. Mountain range at sunrise. Extreme wide shot, slow crane rising to reveal the full valley. Golden hour lighting, volumetric light through the fog, 24mm.
SFX: wind gusting across the ridge, a single crow call.
Ambient: distant birdsong, a faint river below.
Music: sweeping cinematic strings.
(16:9, 1080p, 10s)Best for: openers and travel content. Why it works: a single steady crane keeps an epic reveal coherent where a swooping move would wobble.
4. Two-person dialogue
Two friends sit across a small table, one leaning in as the other reacts. Warm bistro interior, afternoon light through a large window. Over-the-shoulder shot, static locked-off camera. High-key lighting, softbox fill, shallow depth of field, 50mm.
Dialogue: she says, "You actually did it — you quit today?" He grins, "This morning."
SFX: cutlery clinking, a chair scraping nearby.
Ambient: low bistro chatter, warm background music.
(16:9, 1080p, 8s)Best for: character beats and skits. Why it works: two short lines fit the clip, and a locked-off camera lets Wan spend its motion budget on synced lips.
5. Image-to-video product turntable
[Start frame: a clean product photo of a running shoe on a white background]
Add motion: the shoe rotates slowly clockwise on an invisible turntable while the studio reflection shifts across its surface. Keep the shoe, colors, and logo exactly as in the source image. Medium close-up, static camera. Preserve the clean white background and soft shadow.
SFX: a subtle low hum.
Music: minimal upbeat electronic.
(1:1, 1080p, 6s)Best for: turning one catalog photo into a rotating clip with accurate branding. Why it works: anchoring "keep the logo exactly as in the source" stops Wan from redrawing the product while it adds motion.
Want more finished examples before you start? Browse the 35 best Wan prompts, read how to prompt Wan for realistic video, or grab ready-made skeletons from Wan prompt templates.
Frequently Asked Questions
What is the Wan prompt formula?
Write six beats in order: Subject, Scene (place plus time), Motion (one primary motion), Lighting, Camera (shot size plus one move), and Style — then add audio lines beneath. Stack the first beats into one flowing paragraph, like a shot brief handed to a crew, and the audio lines carry Wan's native sound.
Does Wan generate sound?
Yes. Wan 2.5 added native audio and Wan 2.7 keeps it: write Dialogue, Voiceover, SFX, Ambient, and Music lines beneath the visual paragraph and the model generates them synced to the picture. In image-to-video you can also pass a real audio file so Wan matches lip and body motion to the track.
How many camera moves should one Wan shot have?
One. Wan follows a motion hierarchy, so a single clear primary motion plus one camera move gives clean results. Stacking several equal moves — a pan that also dollies and orbits — makes the shot chaotic. Need a second move? Make it a second clip or use Wan's multi-shot mode with labeled scenes.
What resolution and length does Wan support?
Wan 2.7 outputs up to 1080p at 24fps, with clips generally in the 5-to-15-second range depending on the interface. Resolution, aspect ratio, duration, seed, and negative prompt are chosen as settings in the app or API — they are not written inline in the prompt text.
Can I set aspect ratio or length inside the prompt?
No. Wan has no inline parameters like Midjourney's --ar. The settings note in parentheses at the end of each prompt — such as (16:9, 1080p, 10s) — is a reminder of what to pick in the app or API, not text to paste into the prompt box.
How does Wan image-to-video work?
Upload a start frame and describe only the motion, camera behavior, and mood you want added to it. Wan 2.7 also accepts an optional end frame, so you can define both the first and last image and let the model generate the motion between them. Keep the described change small and physical to stay on-model.
Is Wan free to use commercially?
The Wan model weights are released under Apache 2.0, which permits commercial use without licensing fees. If you run it through a hosted service such as Alibaba Cloud Model Studio or a third-party API, that platform's pricing and terms still apply, but the output itself is yours to use commercially.
What is prompt expansion or thinking mode in Wan?
Prompt expansion (sometimes called thinking mode) lets Wan rewrite a short prompt into a richer, more detailed brief before generating. It is useful for filling in lighting and camera detail you left out, but for tight control over a specific shot, write the full six-beat prompt yourself and leave expansion off.