This is the one-page reference for prompting Meta Muse Video — Meta Superintelligence Labs' text-to-video model, announced July 7, 2026 alongside Muse Image and ranked #3 for text-to-video on the LMArena human-preference leaderboard at preview. It runs across the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp, and its headline edge is native synchronized audio — dialogue, ambience, sound effects and music generated in the same pass as the picture. Like Muse Image, it takes full natural-language prose: no comma tags, no --flags, no numeric parameters.

The prose formula, in one line: subject + appearance → one clear action → explicit camera move → setting → lighting & mood → audio line (name the source of each sound) → aspect ratio in words. One continuous action and one camera move per shot; for sequences, write Shot 1… Shot 2… Shot 3… beats. New to it? Start with the best Meta Muse Video prompts roundup, then keep this sheet open while you build. For the full walkthrough, see how to prompt Meta Muse Video for realistic video.

Advertisement

Camera moves

Muse follows camera direction closely. Name one move per shot and drop the example phrase straight into the camera slot of the formula.

MoveWhat it doesExample phrase
Push-inCamera moves toward the subject — builds focus and tension"a slow push-in on her face"
Pull-backCamera retreats — reveals context and scale"a smooth pull-back to reveal the whole valley"
DollyCamera glides on a straight line, level and steady"a steady dolly forward down the corridor"
TrackingCamera moves alongside a moving subject, matching pace"a tracking side shot keeping pace with the runner"
OrbitCamera arcs around the subject — adds depth and drama"a slow orbit around the two of them"
AerialHigh overhead or drone move — scale and geography"an aerial drone push over the treetops"
Crane upCamera rises vertically — a lifting, epic reveal"a crane up from the street to the rooftops"
TiltCamera pivots up or down from a fixed point"a slow tilt up from her boots to her face"
PanCamera pivots left or right, sweeping across a scene"a gentle pan across the skyline"
Whip panFast blurred pan — a snappy transition or reaction"a whip pan to the door as it slams"
Static lock-offCamera holds perfectly still — motion happens in-frame"a static lock-off as steam rises"

Shot sizes & framing

The shot size decides how close the viewer sits. State it in words at the start of the visual description.

FramingWhat it doesExample phrase
Extreme wideTiny subject in a vast environment — sets place and scale"an extreme wide shot of a lone hiker on a ridge"
WideFull subject with surroundings — establishes the scene"a wide shot of the market street"
MediumWaist-up — the everyday, conversational frame"a medium shot of the chef at the counter"
Close-upFace or object fills the frame — emotion and detail"a close-up of her hands kneading dough"
Extreme close-upA single detail — an eye, a droplet, a texture"an extreme close-up of a droplet on a petal"
Over-the-shoulderFramed past one person's shoulder — dialogue and POV context"an over-the-shoulder shot facing the stranger"
POVThe camera is the subject's eyes — immersive first person"a POV shot walking through the neon alley"

Audio layers

Sound is generated in the same pass as the picture, so describe the sound world and, for each sound, name the visible object that makes it. Layer these four freely; keep spoken lines short since sync is still improving.

LayerHow to phrase it — name the visible sourceExample phrase
DialoguePut the line in quotes and say who speaks it; keep it short"the old sailor says, "Storm's coming in""
AmbienceName the space and its background wash of sound"soft ambient chatter and clinking cups fill the cafe"
Sound effects / foleyTie each effect to the object on screen making it"waves crashing on the rocks, gulls calling overhead"
Music cueDescribe the mood, tempo and instrument of the score"a slow, warm piano cue rising under the shot"

Aspect ratios in words

State the ratio in words at the end of the prompt — Muse Video takes no --ar flags and no numeric parameters. Describe length in words too ("one short continuous shot"), never a number of seconds.

Ratio (in words)Best forHow to end the prompt
Square 1:1Feed posts, logo loops, product tiles"Square 1:1, one continuous shot."
Landscape 16:9YouTube, cinematic b-roll, desktop"Widescreen 16:9, one smooth shot."
Vertical 9:16Reels, Stories, TikTok, phone-first"Vertical 9:16 format, a single continuous shot."
Portrait 4:5Instagram feed portrait — max vertical the feed allows"Portrait 4:5, one unbroken take."
Classic 4:3Retro, archival and documentary looks"Classic 4:3, a single steady shot."
Advertisement

Modes

Muse Video runs in three modes. Pick by what you already have, and add the key line so the model knows the job.

ModeWhen to useThe key line
Text-to-video (T2V)You want Muse to invent the whole scene from a descriptionWrite the full prose prompt — subject, action, camera, setting, lighting, audio, ratio.
Image-to-video (I2V)You already have the exact frame and only want it to move"Animate this image: [motion + audio to add]. Keep the face, wardrobe and background identical."
Video-to-video (V2V)You want to restyle or extend an existing clip"Restyle this clip as [new look], keeping the original motion" — or "continue the action: [what happens next]."

Style & mood presets

The style word sets the whole grade. Name one clear look in the style slot; stacking three fights the model.

Style / moodResultExample phrase
CinematicFilm colour grade, shallow focus, atmospheric haze"a cinematic look with a warm filmic grade"
DocumentaryNatural, handheld, unstyled realism"a candid documentary feel, natural light"
Film-noirHigh-contrast black-and-white, deep shadows, tension"a moody film-noir look, hard shadows"
Golden-hourWarm low sun, long shadows, glowing edges"bathed in soft golden-hour light"
MoodyLow-key, desaturated, quiet and heavy"a moody, desaturated atmosphere"
VibrantPunchy saturated colour, high energy"a vibrant, high-saturation commercial look"
AnimeClean line art, cel shading, stylized motion"a clean anime style with crisp line art"
ClaymationStop-motion clay with fingerprint texture"a claymation stop-motion look"
Retro filmGrain, faded colour, vintage tone"a grainy retro-film look, faded colours"

Lighting cues

Lighting sets mood faster than anything else. Name a real setup and Muse renders it; drop one phrase into the lighting slot.

CueLookExample phrase
Soft window lightGentle directional daylight — flattering for people and food"lit by soft window light from the left"
Golden-hour backlightWarm halo and long shadows — dreamy and cinematic"warm golden-hour backlight rimming the subject"
Rim lightBright edge separating subject from a dark background"a crisp rim light against the shadows"
Low-key lightMostly dark with selective highlights — noir and tension"low-key lighting, deep pools of shadow"
High-key lightBright, airy, low-shadow — clean and optimistic"bright high-key light, airy and clean"
Neon practical lightsColoured urban glow on wet surfaces — magenta and cyan"neon practical lights reflecting on wet asphalt"
Hard midday sunCrisp high-contrast shadows — bold and graphic"hard midday sun with sharp shadows"
Candlelight / firelightWarm, flickering, low and intimate"warm flickering firelight on their faces"
MoonlightCool blue night light — quiet and atmospheric"cool blue moonlight through the window"
Advertisement

Example prompts

Eight complete prompts that combine the tables above. Each is one flowing prose paragraph that ends with the ratio in words. Paste any into the Meta AI app or meta.ai and swap the bracketed parts. For more, browse the best Meta Muse Video prompts.

Reminder: Muse Video takes no --flags and no numeric parameters — everything is prose, and the aspect ratio is stated in words at the end. Never write a Settings: line with px, fps or seconds.

1. Cinematic valley reveal (T2V)

An extreme wide shot of a misty mountain valley at dawn, layers of pine ridges fading into low fog as thin clouds drift between the peaks. A slow aerial drone push forward glides over the treetops toward the rising sun. Soft golden-hour backlight breaks on the horizon, a cinematic warm grade, calm and epic. Wind hushing through the pines with a slow, warm string cue rising underneath. Widescreen 16:9, one smooth shot.

Best for: title backgrounds and establishing b-roll. One aerial move plus drifting fog gives Muse a single clear direction of motion.

2. Barista close-up with ambience (T2V)

A close-up of a barista's hands pouring steamed milk into a cup to form a leaf of latte art, steam curling upward. A slow push-in tightens on the cup and her hands. Soft window light from the left with a warm bounce, shallow focus, a cosy documentary feel. Soft ambient chatter and clinking cups fill the cafe, milk steamer hissing on the machine. Portrait 4:5, one unbroken take.

Why it works: a contained hand motion plus a slow push-in keeps detail sharp, and the named sources trigger clean synced ambience.

3. Talking-head with short dialogue (T2V)

A medium shot of a weathered fisherman in a knitted grey sweater standing on a harbour wall at dusk, looking just off camera. A static lock-off holds steady as the wind lifts his collar. Warm golden-hour backlight rims his shoulders, a cinematic grade, quiet and reflective. He says, "Storm's coming in," his voice low and gravelly, gulls calling overhead and waves slapping the wall. Widescreen 16:9, one continuous shot.

Best for: dialogue beats. Keep the spoken line to a clause or two so sync holds.

4. Vertical product hero (T2V)

A close-up of a faceted glass perfume bottle with a gold cap on a wet black reflective surface, a soft swirl of pink mist drifting behind it. A slow orbit arcs around the bottle. Elegant studio lighting with a bright rim highlight on the edges, a vibrant high-saturation commercial look, crisp reflections. A soft airy synth cue and a gentle glassy chime as the mist settles. Vertical 9:16 format, a single continuous shot.

Why it works: a single-axis orbit on a clean set is the classic product-hero move, and drifting mist adds life without warping the glass.

5. Anime rooftop sunset (T2V)

A medium shot of a young anime girl with short teal hair standing on a school rooftop at sunset, her hair and skirt swaying in the wind as she turns her head toward the camera and smiles softly. A slow push-in settles on her face. Warm orange light and long shadows, a clean anime style with crisp line art and a detailed sky. Gentle wind, distant city hum, and a wistful piano cue. Widescreen 16:9, one continuous shot.

Best for: painterly openers. Swaying hair supplies motion while her face stays the calm anchor.

6. Action beat with VFX (T2V)

A wide low-angle shot of a caped hero landing hard on a rain-slicked rooftop at night, cracking the concrete with a shockwave of dust and water as the cape settles. A slow push-in from below holds on the impact. Cool blue moonlight with a warm city glow behind, a cinematic blockbuster grade. A deep bass boom on impact, debris skittering, rain hissing, and a tense low drone under it. Widescreen 16:9, one smooth shot.

Why it works: the impact gives a burst of motion up front, then the push-in lets it settle into a hero beat.

7. Image-to-video — living portrait (I2V)

Animate this portrait: the person blinks naturally and gives a soft warm smile, hair moving slightly in a gentle breeze, a subtle head turn toward the camera. A slow subtle push-in adds depth. Keep the face, wardrobe, lighting and background identical to the photo. Quiet room tone with a faint, warm ambient pad underneath. Portrait 4:5, one unbroken take.

Best for: bringing a still to life. Naming small, natural motions plus a keep-identical line avoids distorting the face.

8. Three-shot sequence (T2V)

A cinematic clip at golden hour, warm grade. Shot 1: a wide shot of a boy standing on an empty train platform holding a letter, wind lifting his hair, distant rail ambience. Shot 2: a close-up of his determined face as he looks up, a slow push-in, a soft piano cue entering. Shot 3: a tracking shot follows him running toward the departing train, cherry petals swirling, the music swelling and the train horn sounding. Smooth cuts between shots. Widescreen 16:9.

Why it works: Muse holds character and scene consistency across Shot 1 / Shot 2 / Shot 3, so you get a tiny storyboard in one take. Keep each shot to one clear action.

Frequently Asked Questions

What is Meta Muse Video and who makes it?

Meta Muse Video is a text-to-video model from Meta Superintelligence Labs (MSL), announced on July 7, 2026 alongside Meta Muse Image and ranked #3 for text-to-video on the LMArena human-preference leaderboard at preview. It is built on the same pretraining base as Muse Image and is rolling out to creators through the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp. Its headline feature is native synchronized audio — dialogue, ambience, sound effects and music are generated in the same pass as the picture.

How do I write a good Meta Muse Video prompt?

Write one flowing prose paragraph, not comma tags and not --flags. Follow the order subject and appearance, one clear action or motion, an explicit camera move, the setting, lighting and mood, an audio line, and finish by stating the aspect ratio in words. Keep it to one continuous action and one camera move per shot, describe length in words like a single continuous take rather than a number of seconds, and use Shot 1, Shot 2, Shot 3 beats for sequences. Muse reasons and self-revises before finalizing, so long, detailed, well-structured prompts pay off.

How does native synchronized audio work in Muse Video?

Sound is generated in the same pass as the picture, so you describe the whole sound world inside the prompt. Name which visible object makes each sound — for example, waves crashing on the rocks, a kettle whistling on the stove, tyres hissing on wet asphalt — and Muse syncs it to the on-screen action. You can layer dialogue in quotes, background ambience, specific sound effects and a music cue. Meta notes audio–video sync is still improving, so keep spoken lines short and do not rely on perfect lip-sync.

What is the difference between text-to-video, image-to-video and video-to-video?

Text-to-video (T2V) builds the whole clip from your written description. Image-to-video (I2V) animates a still you upload — describe only the motion and audio to add and lock identity with a line like keep the face, wardrobe and background identical. Video-to-video (V2V) restyles or extends an existing clip, so you state the new look or what happens next while keeping the original motion. Use T2V to invent a scene, I2V when you already have the exact frame, and V2V to transform footage you already have.

Does Meta Muse Video do lip-sync and dialogue?

It can generate spoken dialogue as part of its native synchronized audio — put the line in quotes and name who says it. Meta has said audio–video sync is still improving, so keep spoken lines short, favour one or two clauses, and do not promise frame-perfect lip-sync. For talking-head shots, a close-up plus a short line reads best; long monologues are where sync drifts.

What aspect ratios can Meta Muse Video use and how do I set them?

State the ratio in words at the end of the prompt — there are no --ar flags and no numeric parameters. The options are square 1:1, landscape 16:9, vertical 9:16, portrait 4:5, and classic 4:3. For example end with vertical 9:16 format, a single continuous shot for Reels and TikTok, or widescreen 16:9, one smooth shot for YouTube and cinematic b-roll. Describe length in words too, such as one short continuous shot, rather than a number of seconds.

Where can I use Meta Muse Video?

Meta Muse Video is rolling out to creators through the Meta AI app, on meta.ai in the browser, inside Instagram Reels and Stories, and in WhatsApp. It accepts natural-language prose prompts in all of them. Because you state the aspect ratio and length in words, the same prompt works everywhere, even where there is no ratio toggle in the interface.

How is Meta Muse Video different from Meta Muse Image?

They share the same pretraining base and the same prose-first prompting style — natural-language sentences, no --flags, no numeric params, and the aspect ratio stated in words. The difference is that Muse Video adds motion, an explicit camera move, and native synchronized audio, so a video prompt has extra parts a still does not: one clear action, a camera direction, and an audio line naming the source of each sound. If you know the Muse Image formula, you already know most of the Muse Video one.

Advertisement