This is the one-page reference for prompting Meta Muse Video — Meta Superintelligence Labs' text-to-video model, announced July 7, 2026 alongside Muse Image and ranked #3 for text-to-video on the LMArena human-preference leaderboard at preview. It runs across the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp, and its headline edge is native synchronized audio — dialogue, ambience, sound effects and music generated in the same pass as the picture. Like Muse Image, it takes full natural-language prose: no comma tags, no --flags, no numeric parameters.
The prose formula, in one line: subject + appearance → one clear action → explicit camera move → setting → lighting & mood → audio line (name the source of each sound) → aspect ratio in words. One continuous action and one camera move per shot; for sequences, write Shot 1… Shot 2… Shot 3… beats. New to it? Start with the best Meta Muse Video prompts roundup, then keep this sheet open while you build. For the full walkthrough, see how to prompt Meta Muse Video for realistic video.
Camera moves
Muse follows camera direction closely. Name one move per shot and drop the example phrase straight into the camera slot of the formula.
| Move | What it does | Example phrase |
|---|---|---|
| Push-in | Camera moves toward the subject — builds focus and tension | "a slow push-in on her face" |
| Pull-back | Camera retreats — reveals context and scale | "a smooth pull-back to reveal the whole valley" |
| Dolly | Camera glides on a straight line, level and steady | "a steady dolly forward down the corridor" |
| Tracking | Camera moves alongside a moving subject, matching pace | "a tracking side shot keeping pace with the runner" |
| Orbit | Camera arcs around the subject — adds depth and drama | "a slow orbit around the two of them" |
| Aerial | High overhead or drone move — scale and geography | "an aerial drone push over the treetops" |
| Crane up | Camera rises vertically — a lifting, epic reveal | "a crane up from the street to the rooftops" |
| Tilt | Camera pivots up or down from a fixed point | "a slow tilt up from her boots to her face" |
| Pan | Camera pivots left or right, sweeping across a scene | "a gentle pan across the skyline" |
| Whip pan | Fast blurred pan — a snappy transition or reaction | "a whip pan to the door as it slams" |
| Static lock-off | Camera holds perfectly still — motion happens in-frame | "a static lock-off as steam rises" |
Shot sizes & framing
The shot size decides how close the viewer sits. State it in words at the start of the visual description.
| Framing | What it does | Example phrase |
|---|---|---|
| Extreme wide | Tiny subject in a vast environment — sets place and scale | "an extreme wide shot of a lone hiker on a ridge" |
| Wide | Full subject with surroundings — establishes the scene | "a wide shot of the market street" |
| Medium | Waist-up — the everyday, conversational frame | "a medium shot of the chef at the counter" |
| Close-up | Face or object fills the frame — emotion and detail | "a close-up of her hands kneading dough" |
| Extreme close-up | A single detail — an eye, a droplet, a texture | "an extreme close-up of a droplet on a petal" |
| Over-the-shoulder | Framed past one person's shoulder — dialogue and POV context | "an over-the-shoulder shot facing the stranger" |
| POV | The camera is the subject's eyes — immersive first person | "a POV shot walking through the neon alley" |
Audio layers
Sound is generated in the same pass as the picture, so describe the sound world and, for each sound, name the visible object that makes it. Layer these four freely; keep spoken lines short since sync is still improving.
| Layer | How to phrase it — name the visible source | Example phrase |
|---|---|---|
| Dialogue | Put the line in quotes and say who speaks it; keep it short | "the old sailor says, "Storm's coming in"" |
| Ambience | Name the space and its background wash of sound | "soft ambient chatter and clinking cups fill the cafe" |
| Sound effects / foley | Tie each effect to the object on screen making it | "waves crashing on the rocks, gulls calling overhead" |
| Music cue | Describe the mood, tempo and instrument of the score | "a slow, warm piano cue rising under the shot" |
Aspect ratios in words
State the ratio in words at the end of the prompt — Muse Video takes no --ar flags and no numeric parameters. Describe length in words too ("one short continuous shot"), never a number of seconds.
| Ratio (in words) | Best for | How to end the prompt |
|---|---|---|
| Square 1:1 | Feed posts, logo loops, product tiles | "Square 1:1, one continuous shot." |
| Landscape 16:9 | YouTube, cinematic b-roll, desktop | "Widescreen 16:9, one smooth shot." |
| Vertical 9:16 | Reels, Stories, TikTok, phone-first | "Vertical 9:16 format, a single continuous shot." |
| Portrait 4:5 | Instagram feed portrait — max vertical the feed allows | "Portrait 4:5, one unbroken take." |
| Classic 4:3 | Retro, archival and documentary looks | "Classic 4:3, a single steady shot." |
Modes
Muse Video runs in three modes. Pick by what you already have, and add the key line so the model knows the job.
| Mode | When to use | The key line |
|---|---|---|
| Text-to-video (T2V) | You want Muse to invent the whole scene from a description | Write the full prose prompt — subject, action, camera, setting, lighting, audio, ratio. |
| Image-to-video (I2V) | You already have the exact frame and only want it to move | "Animate this image: [motion + audio to add]. Keep the face, wardrobe and background identical." |
| Video-to-video (V2V) | You want to restyle or extend an existing clip | "Restyle this clip as [new look], keeping the original motion" — or "continue the action: [what happens next]." |
Style & mood presets
The style word sets the whole grade. Name one clear look in the style slot; stacking three fights the model.
| Style / mood | Result | Example phrase |
|---|---|---|
| Cinematic | Film colour grade, shallow focus, atmospheric haze | "a cinematic look with a warm filmic grade" |
| Documentary | Natural, handheld, unstyled realism | "a candid documentary feel, natural light" |
| Film-noir | High-contrast black-and-white, deep shadows, tension | "a moody film-noir look, hard shadows" |
| Golden-hour | Warm low sun, long shadows, glowing edges | "bathed in soft golden-hour light" |
| Moody | Low-key, desaturated, quiet and heavy | "a moody, desaturated atmosphere" |
| Vibrant | Punchy saturated colour, high energy | "a vibrant, high-saturation commercial look" |
| Anime | Clean line art, cel shading, stylized motion | "a clean anime style with crisp line art" |
| Claymation | Stop-motion clay with fingerprint texture | "a claymation stop-motion look" |
| Retro film | Grain, faded colour, vintage tone | "a grainy retro-film look, faded colours" |
Lighting cues
Lighting sets mood faster than anything else. Name a real setup and Muse renders it; drop one phrase into the lighting slot.
| Cue | Look | Example phrase |
|---|---|---|
| Soft window light | Gentle directional daylight — flattering for people and food | "lit by soft window light from the left" |
| Golden-hour backlight | Warm halo and long shadows — dreamy and cinematic | "warm golden-hour backlight rimming the subject" |
| Rim light | Bright edge separating subject from a dark background | "a crisp rim light against the shadows" |
| Low-key light | Mostly dark with selective highlights — noir and tension | "low-key lighting, deep pools of shadow" |
| High-key light | Bright, airy, low-shadow — clean and optimistic | "bright high-key light, airy and clean" |
| Neon practical lights | Coloured urban glow on wet surfaces — magenta and cyan | "neon practical lights reflecting on wet asphalt" |
| Hard midday sun | Crisp high-contrast shadows — bold and graphic | "hard midday sun with sharp shadows" |
| Candlelight / firelight | Warm, flickering, low and intimate | "warm flickering firelight on their faces" |
| Moonlight | Cool blue night light — quiet and atmospheric | "cool blue moonlight through the window" |
Example prompts
Eight complete prompts that combine the tables above. Each is one flowing prose paragraph that ends with the ratio in words. Paste any into the Meta AI app or meta.ai and swap the bracketed parts. For more, browse the best Meta Muse Video prompts.
Reminder: Muse Video takes no --flags and no numeric parameters — everything is prose, and the aspect ratio is stated in words at the end. Never write a Settings: line with px, fps or seconds.
1. Cinematic valley reveal (T2V)
An extreme wide shot of a misty mountain valley at dawn, layers of pine ridges fading into low fog as thin clouds drift between the peaks. A slow aerial drone push forward glides over the treetops toward the rising sun. Soft golden-hour backlight breaks on the horizon, a cinematic warm grade, calm and epic. Wind hushing through the pines with a slow, warm string cue rising underneath. Widescreen 16:9, one smooth shot.Best for: title backgrounds and establishing b-roll. One aerial move plus drifting fog gives Muse a single clear direction of motion.
2. Barista close-up with ambience (T2V)
A close-up of a barista's hands pouring steamed milk into a cup to form a leaf of latte art, steam curling upward. A slow push-in tightens on the cup and her hands. Soft window light from the left with a warm bounce, shallow focus, a cosy documentary feel. Soft ambient chatter and clinking cups fill the cafe, milk steamer hissing on the machine. Portrait 4:5, one unbroken take.Why it works: a contained hand motion plus a slow push-in keeps detail sharp, and the named sources trigger clean synced ambience.
3. Talking-head with short dialogue (T2V)
A medium shot of a weathered fisherman in a knitted grey sweater standing on a harbour wall at dusk, looking just off camera. A static lock-off holds steady as the wind lifts his collar. Warm golden-hour backlight rims his shoulders, a cinematic grade, quiet and reflective. He says, "Storm's coming in," his voice low and gravelly, gulls calling overhead and waves slapping the wall. Widescreen 16:9, one continuous shot.Best for: dialogue beats. Keep the spoken line to a clause or two so sync holds.
4. Vertical product hero (T2V)
A close-up of a faceted glass perfume bottle with a gold cap on a wet black reflective surface, a soft swirl of pink mist drifting behind it. A slow orbit arcs around the bottle. Elegant studio lighting with a bright rim highlight on the edges, a vibrant high-saturation commercial look, crisp reflections. A soft airy synth cue and a gentle glassy chime as the mist settles. Vertical 9:16 format, a single continuous shot.Why it works: a single-axis orbit on a clean set is the classic product-hero move, and drifting mist adds life without warping the glass.
5. Anime rooftop sunset (T2V)
A medium shot of a young anime girl with short teal hair standing on a school rooftop at sunset, her hair and skirt swaying in the wind as she turns her head toward the camera and smiles softly. A slow push-in settles on her face. Warm orange light and long shadows, a clean anime style with crisp line art and a detailed sky. Gentle wind, distant city hum, and a wistful piano cue. Widescreen 16:9, one continuous shot.Best for: painterly openers. Swaying hair supplies motion while her face stays the calm anchor.
6. Action beat with VFX (T2V)
A wide low-angle shot of a caped hero landing hard on a rain-slicked rooftop at night, cracking the concrete with a shockwave of dust and water as the cape settles. A slow push-in from below holds on the impact. Cool blue moonlight with a warm city glow behind, a cinematic blockbuster grade. A deep bass boom on impact, debris skittering, rain hissing, and a tense low drone under it. Widescreen 16:9, one smooth shot.Why it works: the impact gives a burst of motion up front, then the push-in lets it settle into a hero beat.
7. Image-to-video — living portrait (I2V)
Animate this portrait: the person blinks naturally and gives a soft warm smile, hair moving slightly in a gentle breeze, a subtle head turn toward the camera. A slow subtle push-in adds depth. Keep the face, wardrobe, lighting and background identical to the photo. Quiet room tone with a faint, warm ambient pad underneath. Portrait 4:5, one unbroken take.Best for: bringing a still to life. Naming small, natural motions plus a keep-identical line avoids distorting the face.
8. Three-shot sequence (T2V)
A cinematic clip at golden hour, warm grade. Shot 1: a wide shot of a boy standing on an empty train platform holding a letter, wind lifting his hair, distant rail ambience. Shot 2: a close-up of his determined face as he looks up, a slow push-in, a soft piano cue entering. Shot 3: a tracking shot follows him running toward the departing train, cherry petals swirling, the music swelling and the train horn sounding. Smooth cuts between shots. Widescreen 16:9.Why it works: Muse holds character and scene consistency across Shot 1 / Shot 2 / Shot 3, so you get a tiny storyboard in one take. Keep each shot to one clear action.
Frequently Asked Questions
What is Meta Muse Video and who makes it?
Meta Muse Video is a text-to-video model from Meta Superintelligence Labs (MSL), announced on July 7, 2026 alongside Meta Muse Image and ranked #3 for text-to-video on the LMArena human-preference leaderboard at preview. It is built on the same pretraining base as Muse Image and is rolling out to creators through the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp. Its headline feature is native synchronized audio — dialogue, ambience, sound effects and music are generated in the same pass as the picture.
How do I write a good Meta Muse Video prompt?
Write one flowing prose paragraph, not comma tags and not --flags. Follow the order subject and appearance, one clear action or motion, an explicit camera move, the setting, lighting and mood, an audio line, and finish by stating the aspect ratio in words. Keep it to one continuous action and one camera move per shot, describe length in words like a single continuous take rather than a number of seconds, and use Shot 1, Shot 2, Shot 3 beats for sequences. Muse reasons and self-revises before finalizing, so long, detailed, well-structured prompts pay off.
How does native synchronized audio work in Muse Video?
Sound is generated in the same pass as the picture, so you describe the whole sound world inside the prompt. Name which visible object makes each sound — for example, waves crashing on the rocks, a kettle whistling on the stove, tyres hissing on wet asphalt — and Muse syncs it to the on-screen action. You can layer dialogue in quotes, background ambience, specific sound effects and a music cue. Meta notes audio–video sync is still improving, so keep spoken lines short and do not rely on perfect lip-sync.
What is the difference between text-to-video, image-to-video and video-to-video?
Text-to-video (T2V) builds the whole clip from your written description. Image-to-video (I2V) animates a still you upload — describe only the motion and audio to add and lock identity with a line like keep the face, wardrobe and background identical. Video-to-video (V2V) restyles or extends an existing clip, so you state the new look or what happens next while keeping the original motion. Use T2V to invent a scene, I2V when you already have the exact frame, and V2V to transform footage you already have.
Does Meta Muse Video do lip-sync and dialogue?
It can generate spoken dialogue as part of its native synchronized audio — put the line in quotes and name who says it. Meta has said audio–video sync is still improving, so keep spoken lines short, favour one or two clauses, and do not promise frame-perfect lip-sync. For talking-head shots, a close-up plus a short line reads best; long monologues are where sync drifts.
What aspect ratios can Meta Muse Video use and how do I set them?
State the ratio in words at the end of the prompt — there are no --ar flags and no numeric parameters. The options are square 1:1, landscape 16:9, vertical 9:16, portrait 4:5, and classic 4:3. For example end with vertical 9:16 format, a single continuous shot for Reels and TikTok, or widescreen 16:9, one smooth shot for YouTube and cinematic b-roll. Describe length in words too, such as one short continuous shot, rather than a number of seconds.
Where can I use Meta Muse Video?
Meta Muse Video is rolling out to creators through the Meta AI app, on meta.ai in the browser, inside Instagram Reels and Stories, and in WhatsApp. It accepts natural-language prose prompts in all of them. Because you state the aspect ratio and length in words, the same prompt works everywhere, even where there is no ratio toggle in the interface.
How is Meta Muse Video different from Meta Muse Image?
They share the same pretraining base and the same prose-first prompting style — natural-language sentences, no --flags, no numeric params, and the aspect ratio stated in words. The difference is that Muse Video adds motion, an explicit camera move, and native synchronized audio, so a video prompt has extra parts a still does not: one clear action, a camera direction, and an audio line naming the source of each sound. If you know the Muse Image formula, you already know most of the Muse Video one.