These are 18 reusable Meta Muse Video prompt skeletons — fill-in-the-blank templates you copy, edit, and paste into the Meta AI app, meta.ai, or the Muse Video surface. Every one is a single natural-language paragraph, not comma tags and not --flags, because Muse Video reasons over your prompt and self-revises before finalizing, so it rewards detailed, well-structured prose.

To use one, swap each [SQUARE-BRACKET slot] for your own detail, delete any phrase you don't need, and keep the shape intact: one clear action, one explicit camera move, a described setting and mood, an [AUDIO] line that names which visible object makes each sound (put any spoken line in "quotes" and keep it short), and the aspect ratio stated in words at the very end. There is no Settings line and no numbers — describe length in words too. Want finished examples? Start with the best Meta Muse Video prompts roundup, keep the cheat sheet open while you fill these in, and read the how-to-prompt guide for the formula behind the b-roll slots.

Advertisement

Cinematic & b-roll templates

Four skeletons for establishing shots and atmosphere. Pick one camera move per shot, keep the action simple, and let the [AUDIO] slot carry the ambience. See the cinematic b-roll pack for finished versions.

1. Establishing Aerial Reveal Template

A sweeping establishing shot of [LOCATION + KEY LANDMARK], with [ONE MOVING ELEMENT, e.g. mist drifting between the ridges]. The camera performs [CAMERA MOVE: a slow aerial push forward / a crane up and pull-back to reveal the full vista]. Bathe the scene in [LIGHTING/MOOD, e.g. soft golden first light on the horizon], with a cinematic color grade that feels [MOOD, e.g. calm and epic]. Audio: [AMBIENCE, e.g. low wind and distant birdsong], with [SFX naming its source, e.g. a river rushing far below]. Widescreen 16:9, one smooth continuous shot.

How to fill it: make the [LANDMARK] a single anchor and pick just one aerial move so the reveal stays directed. Example: "A sweeping establishing shot of a fjord town under snow-capped peaks, with fog drifting between the ridges; the camera cranes up and pulls back to reveal the full coastline, in soft blue-hour light that feels serene and vast. Audio: low wind, with gulls calling over the water below. Widescreen 16:9, one smooth continuous shot."

2. Slow Push-In Mood Template

An intimate close shot of [SUBJECT/DETAIL, e.g. raindrops running down a café window at night], with [SMALL ONGOING MOTION, e.g. a single drop sliding and merging with others]. The camera holds nearly static with [CAMERA MOVE: a subtle slow push-in]. Light it with [LIGHTING/MOOD, e.g. warm-and-blue tones from blurred street signs beyond the glass], quiet and reflective. Audio: [AMBIENCE + SFX naming source, e.g. soft rain tapping the glass and a low hiss of passing tires]. Vertical 9:16 format, a single unbroken take.

How to fill it: keep the camera almost locked so the small motion reads clean. Example: fill [SUBJECT] with "steam curling from a mug on a windowsill" and [AUDIO] with "a kettle ticking as it cools nearby and faint street noise."

3. Tracking Nature B-Roll Template

A [WIDE / MEDIUM] b-roll shot of [NATURAL SCENE, e.g. a narrow path through an autumn forest carpeted in gold leaves], with [SECONDARY MOTION, e.g. a few leaves drifting down]. The camera performs [CAMERA MOVE: a slow steady tracking move along the path / a gentle side dolly]. Sunlight arrives as [LIGHTING/MOOD, e.g. soft warm shafts through the canopy], peaceful and nostalgic. Audio: [AMBIENCE + SFX, e.g. a soft breeze and rustling leaves, with a single bird call from the trees]. Widescreen 16:9, one continuous shot.

How to fill it: assign the [SECONDARY MOTION] to one element so the frame breathes without the camera fighting it. Example: [NATURAL SCENE] "tall grass on a coastal dune at dusk," [AUDIO] "wind hissing through the grass and waves breaking out of frame."

4. Time-Lapse Cityscape Template

A time-lapse of [CITY SCENE, e.g. a modern skyline shifting from dusk to night], where [CHANGE OVER TIME, e.g. office windows light up one by one and headlight streaks flow on the roads below]. The camera stays calm with [CAMERA MOVE: a very slow push-in toward the tallest tower]. Grade it [LIGHTING/MOOD, e.g. a deep blue-to-orange gradient sky], polished and cinematic. Audio: [AMBIENCE, e.g. a soft rising hum of city traffic, with a distant siren fading in and out]. Widescreen 16:9, one continuous shot.

How to fill it: name the [CHANGE OVER TIME] explicitly — the light shift is what "time-lapse" animates while the camera stays still. Example: [CITY SCENE] "clouds racing over a coastal harbor at sunrise," [AUDIO] "gulls and the clank of rigging from moored boats."

People, talking & dialogue templates

Four skeletons for faces and voices. Each has an [AUDIO: "spoken line"] slot — keep quoted lines short, describe the tone, and treat lip-sync as approximate. See prompts with audio for more.

5. Talking-Head Piece-to-Camera Template

A [SHOT SIZE, e.g. chest-up] shot of [SPEAKER: age, appearance, wardrobe] speaking directly to camera in [SETTING, e.g. a softly lit home office]. They [ONE ACTION, e.g. lean in slightly and gesture once] while delivering the line. The camera holds [CAMERA MOVE: a locked-off frame with a barely perceptible push-in]. Light them with [LIGHTING/MOOD, e.g. soft window light from the left, warm and approachable]. Audio: [SPOKEN LINE in a stated tone, e.g. warm and confident: "Here's the one thing nobody tells you about starting out."], with quiet room tone underneath. Vertical 9:16 format, a single continuous shot.

How to fill it: keep the quoted line to one sentence and name the delivery tone. Example: [SPEAKER] "a woman in her 30s in a mustard sweater," line "Let me show you the fastest way to do this." Expect close but not perfect lip-sync.

6. Two-Person Dialogue Template

A [SHOT SIZE, e.g. two-shot] of [PERSON A] and [PERSON B] in [SETTING, e.g. a busy coffee shop by the window]. [PERSON A] [ONE ACTION, e.g. slides a cup across the table] and speaks, then [PERSON B] reacts. The camera holds [CAMERA MOVE: a slow arc / a static frame favoring both faces]. Light it [LIGHTING/MOOD, e.g. warm afternoon light through the glass], natural and candid. Audio: [PERSON A: "short spoken line"] answered by [PERSON B: "short spoken line"], over [AMBIENCE naming source, e.g. an espresso machine hissing and low café chatter]. Widescreen 16:9, one continuous shot.

How to fill it: give each speaker one short line so the exchange fits a single shot. Example: A: "Did it actually work?" B, grinning: "See for yourself." over "cups clinking and muffled conversation."

7. Candid Reaction Template

A candid [SHOT SIZE, e.g. medium close-up] of [SUBJECT] reacting to [TRIGGER, e.g. opening a surprise gift], caught mid-[EXPRESSION, e.g. a laugh that turns to happy disbelief]. The camera stays [CAMERA MOVE: a loose handheld frame easing in toward their face]. Light it with [LIGHTING/MOOD, e.g. soft natural window light], warm and unposed. Audio: [SPONTANEOUS SOUND, e.g. a gasp then a quiet "no way…"], with [AMBIENCE naming source, e.g. wrapping paper crinkling in their hands]. Portrait 4:5, a single continuous shot.

How to fill it: "candid" plus "mid-[EXPRESSION]" stops the model from planning a stiff, posed beat. Example: [TRIGGER] "tasting the dish for the first time," sound "eyes widening with a soft 'oh, that's good.'"

8. Voiceover Narrator Template

A [SHOT TYPE, e.g. cinematic b-roll] of [SCENE/SUBJECT] with [ONE ACTION OR MOTION], carried by [CAMERA MOVE: a slow push-in / a gentle tracking move]. Set the mood with [LIGHTING/MOOD]. Over the images, a [VOICE DESCRIPTION, e.g. calm, gravelly narrator] delivers the voiceover: [AUDIO: "one or two short sentences of narration"], with soft [MUSIC CUE, e.g. a slow ambient piano bed] and [AMBIENCE naming source] underneath. [ASPECT RATIO IN WORDS, e.g. Widescreen 16:9], one continuous shot.

How to fill it: the narration is off-screen, so there's no lip-sync to worry about — describe the voice and keep the line tight. Example: narrator over a foggy harbour, "Every morning, before the town wakes, the sea does its quiet work."

Advertisement

Product & ad templates

Three skeletons for selling. Center the product, choose one controlled camera move, and use audio to add polish. Keep the identity of the product exact and the background clean.

9. Product Hero Rotation Template

A hero shot of [PRODUCT: name, material, color, finish] on [SURFACE, e.g. a wet black reflective surface], with [ATMOSPHERE, e.g. a soft swirl of mist drifting behind it]. The product performs [ONE ACTION, e.g. a slow 360-degree rotation] as the camera holds [CAMERA MOVE: a locked frame / a slow macro push-in]. Light it with [LIGHTING/MOOD, e.g. elegant studio light with a bright rim highlight on the edges], premium and clean. Audio: [SFX + MUSIC naming feel, e.g. a soft whoosh as it turns over a minimal ambient tone]. Square 1:1, one smooth continuous shot.

How to fill it: name the [material] and [finish] precisely — the model renders highlights and reflections straight from those words. Example: [PRODUCT] "a faceted glass perfume bottle with a gold cap," audio "a delicate chime as light sweeps the facets."

10. Product In-Use Lifestyle Template

A lifestyle shot of [PERSON] using [PRODUCT] in [REAL SETTING where a customer would use it], performing [ONE ACTION, e.g. lifting the mug and taking a sip]. The camera follows with [CAMERA MOVE: a slow push-in on the product / a soft tracking move]. Light it with [LIGHTING/MOOD, e.g. bright natural morning light], warm and aspirational, product in sharp focus. Audio: [AMBIENCE + SFX naming source, e.g. a soft kitchen hum and the gentle clink of the mug on the counter], with an optional [MUSIC CUE]. Portrait 4:5, one continuous shot.

How to fill it: choose props and a setting that imply the buyer, not just decoration. Example: [PRODUCT] "a stainless travel bottle," [SETTING] "a sunlit trailhead," audio "birdsong and the click of the cap opening."

11. Founder Testimonial Ad Template

A [SHOT SIZE, e.g. chest-up] shot of [FOUNDER/CUSTOMER: appearance] speaking to camera in [SETTING, e.g. their workshop], holding or gesturing to [PRODUCT]. They [ONE ACTION, e.g. hold up the product and smile] as the camera holds [CAMERA MOVE: a locked frame with a slight push-in]. Light it [LIGHTING/MOOD, e.g. warm practical light], sincere and grounded. Audio: [SPOKEN LINE in a stated tone, e.g. warm and genuine: "We built this because nothing else did the job."], over quiet room tone and a soft [MUSIC CUE]. Vertical 9:16 format, a single continuous shot.

How to fill it: one short, believable line beats a scripted pitch. Example: [FOUNDER] "a man in an apron in a bakery," line "Every loaf still starts at 4am, by hand."

Social / Reels templates

Three vertical skeletons built for the feed. Lead with a hook, keep it to one action, and always finish with vertical 9:16. See the social media prompts pack.

12. Vertical Hook Reel Template

A punchy vertical clip that opens on [ATTENTION HOOK, e.g. a close-up of the finished result], where [SUBJECT] [ONE ACTION, e.g. flips the pan and the food tosses in the air]. The camera stays [CAMERA MOVE: a tight handheld frame with a quick push-in]. Light it [LIGHTING/MOOD, e.g. bright and high-energy]. Audio: [ENERGETIC SFX naming source, e.g. a sizzle and a satisfying sizzle-pop], with an upbeat [MUSIC CUE], and a spoken hook: [AUDIO: "you're doing this wrong — watch."]. Vertical 9:16 format, a single continuous shot.

How to fill it: put the payoff in the first frame, then the action — that's the scroll-stopper. Example: [HOOK] "a phone screen showing 1M views," line "here's the exact hook I used."

13. How-To Demo Reel Template

A vertical demo of [TASK, e.g. folding a fitted sheet], shown from [ANGLE, e.g. a top-down over-the-shoulder view] as [HANDS/PERSON] perform [ONE CLEAR STEP, e.g. tucking the corners together]. The camera holds [CAMERA MOVE: a steady locked overhead frame]. Light it [LIGHTING/MOOD, e.g. clean, even, bright]. Audio: a [VOICE TONE, e.g. friendly, clear] voice explaining, [AUDIO: "start by matching these two corners"], with soft [SFX naming source, e.g. fabric rustling] underneath. Vertical 9:16 format, a single continuous shot.

How to fill it: one step per clip keeps a how-to legible; chain steps as a sequence if you need more. Example: [TASK] "brewing pour-over coffee," line "pour in slow circles, not straight down."

14. Trend Transition Reel Template

A vertical clip built around a single transition: [BEFORE STATE, e.g. a bare room], then [SUBJECT] performs [ONE ACTION that motivates the cut, e.g. sweeps a hand across the frame], revealing [AFTER STATE, e.g. the fully styled room]. The camera uses [CAMERA MOVE: a whip pan / a quick push-in on the gesture]. Light it [LIGHTING/MOOD, e.g. bright and vivid]. Audio: [SFX naming source, e.g. a swoosh on the whip pan] landing on a beat drop in an upbeat [MUSIC CUE]. Vertical 9:16 format, a single continuous shot.

How to fill it: tie the transition to a physical gesture so the cut feels motivated, not random. Example: [BEFORE] "plain t-shirt," gesture "spins once," [AFTER] "full styled outfit," SFX "a whoosh into a bass hit."

Advertisement

Image-to-video templates

Two skeletons that animate a still you upload. Describe only the motion and audio to add, and keep the identity-lock line intact so the frame stays true. See image-to-video prompts for more.

15. Photo-to-Motion Portrait Template

Animate this photo of [SUBJECT]: they [SMALL NATURAL MOTION, e.g. blink and give a soft warm smile], with [SECONDARY MOTION, e.g. hair moving slightly in a gentle breeze], as the camera adds [CAMERA MOVE: a slow subtle push-in]. Keep the face, wardrobe, lighting and background identical to the photo — change nothing about the composition. Audio: [AMBIENCE + optional short line naming source, e.g. soft outdoor wind, with a quiet "hi there" from the subject]. [ASPECT RATIO IN WORDS matching the source, e.g. Portrait 4:5], a single continuous shot.

How to fill it: ask for small, natural motions only, and always keep the identity-lock sentence so the likeness holds. Example: [SUBJECT] "an older man on a porch," motion "a slow nod and a smile," audio "cicadas and a creaking rocking chair."

16. Product Still to Ad Template

Animate this product still: add [MOTION, e.g. a slow drift of steam rising / a gentle rotation], with [ATMOSPHERE, e.g. soft light sweeping across the surface], as the camera does [CAMERA MOVE: a slow macro push-in]. Keep the product, its label, colors and background identical to the image — do not alter the packaging or composition. Audio: [SFX + MUSIC naming feel, e.g. a soft ambient tone with a subtle sparkle as the light passes]. [ASPECT RATIO IN WORDS matching the source, e.g. Square 1:1], one smooth continuous shot.

How to fill it: lock the label and packaging explicitly — that's what keeps a real product true when it moves. Example: [MOTION] "condensation beading and a droplet sliding down the can," audio "a crisp fizz and a light chime."

Multi-shot sequence templates

Two skeletons that hold character and scene consistency across shots. Write Shot 1… Shot 2… Shot 3… beats and keep each shot to one action and one camera move.

17. Three-Shot Story Sequence Template

[OVERALL STYLE + MOOD, e.g. warm cinematic, golden hour], keeping [CHARACTER] consistent across all shots. Shot 1: a [WIDE] establishing shot of [CHARACTER] in [SETTING], [ONE ACTION], camera [CAMERA MOVE]. Shot 2: a close-up of [CHARACTER'S FACE / DETAIL] as [ONE ACTION/EXPRESSION], camera [CAMERA MOVE]. Shot 3: [CHARACTER] [FINAL ACTION that resolves the beat], camera [CAMERA MOVE]. Smooth cuts between shots. Audio: [AMBIENCE naming source] throughout, with [optional short spoken line in Shot 2, e.g. "almost there."]. Widescreen 16:9, a short continuous sequence.

How to fill it: reuse the same [CHARACTER] description in every shot so the model holds identity. Example: a girl in a red coat waits on a platform (Shot 1), her determined face looks up (Shot 2), she runs toward the train as petals swirl (Shot 3).

18. Three-Shot Product Sequence Template

[OVERALL STYLE, e.g. clean premium commercial], keeping [PRODUCT] and its packaging consistent across all shots. Shot 1: a hero shot of [PRODUCT] on [SURFACE], [ONE ACTION, e.g. a slow rotation], camera [CAMERA MOVE]. Shot 2: a macro close-up of [KEY DETAIL, e.g. the texture / logo], camera [CAMERA MOVE]. Shot 3: [PERSON] using [PRODUCT] in [SETTING], [ONE ACTION], camera [CAMERA MOVE]. Smooth cuts between shots. Audio: [SFX naming source] and a soft [MUSIC CUE] throughout, optionally ending on a spoken tagline: [AUDIO: "short tagline."]. Vertical 9:16 format, a short continuous sequence.

How to fill it: repeat the [PRODUCT] description in each shot and keep every shot to one action. Example: a sneaker rotates on a plinth (Shot 1), macro of the stitching (Shot 2), a runner laces up and strides off (Shot 3), tagline "Built to move."

Once a filled-in template works, save your version and reuse the exact wording — consistent phrasing gives consistent results. When you want polished, ready-to-paste prompts instead of skeletons, go back to the best Meta Muse Video prompts roundup, and keep the cheat sheet nearby for the camera-move and audio rules.

Frequently Asked Questions

How do I use one of these Meta Muse Video templates?

Copy the whole block, then replace every [bracketed slot] with your own detail. Keep the natural-language paragraph shape — Meta Muse Video reads full descriptive sentences, not comma tags or --flags, and its agentic reasoning plans the shot and self-revises from the way you phrase things. Keep one clear action and one camera move per shot, describe the sound world including which visible object makes each sound, and leave the aspect-ratio sentence in words at the end. Paste the finished paragraph into the Meta AI app, meta.ai, or the Muse Video surface and generate.

What is Meta Muse Video and who makes it?

Meta Muse Video is a text-to-video model from Meta Superintelligence Labs, announced on July 7, 2026 alongside Meta Muse Image and built on the same pretraining base. At preview it ranked #3 for text-to-video on the LMArena human-preference leaderboard. Its headline feature is native synchronized audio — dialogue, ambience, sound effects and music generated in the same pass as the picture. It is rolling out to creators through the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp.

Why are these templates written as prose instead of tags or a settings line?

Meta Muse Video reasons over your prompt, uses tools, and self-revises its draft before finalizing, so it rewards long, detailed, well-structured prose and follows complex multi-part instructions faithfully. It does not use comma-tag lists, --flags, or a numeric Settings line. Meta has not published exact resolution, frame rate or duration numbers, so you never write 1080p, 24fps or 8s — you describe length in words like one continuous shot and state the aspect ratio in words at the end.

How does the native audio slot work?

Sound is generated in the same pass as the picture, so every template includes an [AUDIO] slot. Describe the whole sound world — ambience, specific sound effects, any music cue — and name which visible object makes each sound, for example the kettle whistling on the stove. For a spoken line, put the exact words in quotes and keep them short. Meta notes audio-video sync is still improving, so short quoted lines work better than long monologues and you should not expect perfect lip-sync.

Does Meta Muse Video do dialogue and lip-sync?

Yes, it can generate spoken dialogue as part of its native audio, and the people, talking and dialogue templates here include a quoted-line slot. Keep the spoken line short — a sentence or two — describe the speaker's tone, and treat lip-sync as approximate rather than frame-perfect, because Meta says audio-video sync is still improving. For anything longer, break it across a multi-shot sequence with one short line per shot.

What is the difference between text-to-video, image-to-video and video-to-video?

Text-to-video (T2V) builds the whole clip from your written description — most templates here are T2V. Image-to-video (I2V) animates a still you upload: you describe only the motion and audio to add and lock identity with a line like keep the face, wardrobe and background identical. Video-to-video (V2V) restyles or extends an existing clip. The image-to-video templates on this page show the identity-lock pattern; use it whenever you start from a photo.

How do I set the aspect ratio in Meta Muse Video?

State it in words at the end of the prompt, for example Vertical 9:16 format, a single continuous shot. Muse Video reads ratios written out plainly — square 1:1, landscape 16:9, vertical 9:16, portrait 4:5, and classic 4:3 are all reliable. Use vertical 9:16 for Reels and Stories, landscape 16:9 for cinematic and YouTube, square 1:1 or portrait 4:5 for the feed. Every template here ends with that sentence so you just swap the ratio.

How is Meta Muse Video different from Meta Muse Image?

They share the same pretraining base and the same prose-first prompting style, but Muse Image makes stills while Muse Video makes short clips with motion and native synchronized audio. So a video prompt adds two things an image prompt does not: one explicit camera move and one audio line. Everything else — natural-language paragraphs, no numeric parameters, aspect ratio stated in words — carries straight over from the image side.

How long should a Meta Muse Video prompt be?

Each filled-in template lands around 35 to 80 words for a single shot, which is enough to name the subject and appearance, one action, one camera move, the setting, the lighting and mood, and the audio without burying the model. Because Muse Video reasons over structured prose, a well-organized longer paragraph works better than a terse one. For a story, use the multi-shot sequence templates and keep each Shot to one action and one camera move.

Advertisement