Realistic clips from Meta Muse Video come down to how you brief it. Announced by Meta Superintelligence Labs on July 7, 2026 and built on the same pretraining base as Muse Image, it wants a single natural-language paragraph — not tags, not --flags, not a settings line. This guide gives you the prose formula, the camera moves it follows, how to write its native synchronized audio, and 9 complete prompts you can paste right now. For the full library, start with the best Meta Muse Video prompts.

Advertisement

The prose video formula

Write one flowing paragraph in this order: subject and appearance → one clear action → one explicit camera move → setting → lighting and mood → an audio line → the aspect ratio in words. Muse reasons over your prompt and self-revises its draft before rendering, so it rewards detailed, well-structured prose and follows complex instructions faithfully.

Muse Video wants natural-language paragraphs, the same style as Muse Image, because an agentic model reads a prompt like a director's brief rather than a keyword list. Comma-tag soup throws away the relationships between subject, camera, and sound. And you state the ratio and length in words — "vertical 9:16 format, a single continuous shot" — because Meta has not published exact resolution, fps, or duration figures. There is no Settings: line and no numeric parameters here; asking for "1080p", "24fps", or "8s" does nothing. Here is the skeleton:

[SUBJECT + appearance] [ONE clear action or motion]. [ONE explicit camera move]. [SETTING with real detail]. [LIGHTING + mood]. [AUDIO: ambience, SFX naming the source, a short line of dialogue in quotes, or a music cue]. [Aspect ratio in words, a single continuous shot].

The two rules that keep clips clean: one continuous action and one camera move per shot. Give the model a single thing to animate and a single direction to move the frame, and it holds motion, identity, and environment coherent. Here it is filled in:

A weathered fisherman in a yellow raincoat hauls a rope hand over hand at the stern of a small boat. The camera slowly pushes in on his face and hands. Grey open sea under a heavy overcast sky, spray misting the lens. Cold desaturated light, tense and physical mood. Native audio: waves slapping the hull, wind buffeting the mic, the wet creak of the rope, and he grunts once, "Almost there." Widescreen 16:9, one smooth continuous shot.

Why it works: one action (hauling the rope), one move (a slow push-in), and every sound is tied to a visible source, so Muse has a clear brief to reason over instead of competing instructions. Keep the Meta Muse Video prompt cheat sheet open while you build these — it lists the camera moves and the ratio-in-words options on one page.

Camera moves Muse follows

Muse has strong prompt adherence, so it follows real camera language closely — but only if you pick one move per shot. Name a specific move rather than a vague "cinematic camera", and the frame gains direction instead of drift.

The moves it handles cleanly: slow push-in and pull-back, dolly (forward or lateral), tracking shot (following alongside or behind a subject), orbit (arcing around the subject), aerial or drone rise, crane up, tilt (up or down), pan (left or right), whip pan for a fast transition, and a static lock-off when you want the subject's own motion to carry the shot. Pair the move to the action: a push-in for intimacy and tension, an orbit to add depth around a still subject, a tracking shot to follow movement, an aerial for scale.

A lone cyclist rides along a coastal cliff road at golden hour, jacket rippling in the wind. The camera tracks alongside her at the same speed, the sea glittering far below. Warm low backlight rimming her silhouette, long shadows, uplifting and cinematic mood. Native audio: wind rushing past, the steady tick of the freewheel, distant gulls, and faint waves below. Widescreen 16:9, one continuous take.

Why it works: the tracking shot moves with the single subject and her single action, so the motion has one clear direction and the horizon stays stable.

A potter's wet hands shape a spinning clay bowl on a wheel, the walls rising and narrowing under her fingers. The camera slowly orbits from the side around to the front. Warm studio window light, water glistening on the clay, calm and focused mood. Native audio: the low hum of the wheel, the wet slick of clay under her palms, and quiet room tone. Portrait 4:5, a single unbroken shot.

Best for: craft and process b-roll. A slow orbit adds depth while the contained hand motion stays sharp.

Writing native audio into the prompt

Muse Video generates native synchronized audio in the same pass as the picture, so dialogue, ambience, sound effects, and music all come from one prompt. Describe the sound world in plain language and name which visible object makes each sound, so the model can line the audio up with the motion.

Handle each layer deliberately. Keep dialogue short and in "quotes" — one brief line per shot, since Meta notes audio-video sync is still improving and long speech will drift. Set the ambience (room tone, wind, city hum). Name SFX by their source ("the kettle whistling on the stove", not just "a whistle"). Add a music cue only when you want it ("a soft piano motif underneath"). Do not promise perfect lip-sync.

A barista slides a finished latte across a marble counter toward the camera and looks up with a small smile. The camera holds in a static lock-off, then eases into a subtle push-in. A warm morning café, sun through the front window, shallow focus, cozy documentary mood. Native audio: the hiss of the steam wand winding down, cups clinking, low café chatter, and she says brightly, "One oat latte." Vertical 9:16 format, a single continuous shot.

Why it works: one short quoted line plus source-named SFX gives Muse audio it can actually sync to the visible actions, and the near-static camera keeps the moment intimate.

Rain streams down a café window at night while warm bokeh from street signs glows beyond the glass; a single drop slides and merges with others. The camera holds static with a barely perceptible push-in. Moody warm-and-blue tones, quiet and intimate mood. Native audio: steady rain against the glass, the muffled swish of a passing car, distant thunder rolling in, and a soft piano motif underneath. Vertical 9:16 format, one short continuous take.

Best for: mood shots and vertical intros. Naming rain, car, and thunder as separate sources lets Muse layer the soundscape cleanly.

Advertisement

Multi-shot sequences

Muse keeps character and scene consistent across shots, so you can write a tiny storyboard in one prompt using Shot 1 / Shot 2 / Shot 3 beats. Set the look once up front, then give each shot one action and one camera move — the model holds identity, wardrobe, and environment across the cuts.

A cinematic sequence in a rain-soaked neon city at night, moody teal-and-magenta grade, the same young woman in a red coat throughout. Shot 1: she steps out of a doorway and opens a black umbrella, slow push-in on her face. Shot 2: a tracking shot follows behind her down a glowing wet alley. Shot 3: she pauses and looks back over her shoulder, a slow pull-back revealing the empty street. Native audio: steady rain, footsteps splashing, distant traffic, and a low ambient synth pad. Smooth cuts between shots. Widescreen 16:9, one continuous sequence.

Why it works: the shared look and character are stated once, and each shot carries exactly one action and one move, so Muse keeps the woman consistent while the story advances.

A warm sunrise cooking sequence in a bright kitchen, the same chef in a linen apron throughout, natural window light and a soft handheld feel. Shot 1: he cracks an egg into a hot pan, close-up push-in as it sizzles. Shot 2: he flips the omelette with a flick of the wrist, static lock-off. Shot 3: he slides it onto a plate and slides the plate toward camera, a slow pull-back. Native audio: the egg sizzling, the pan scraping, and he says, "Breakfast." Smooth cuts. Vertical 9:16 format, one continuous sequence.

Best for: short social recipes and how-tos where the same person and kitchen must carry across three beats.

Image-to-video & video-to-video

In image-to-video (I2V) you upload a still and Muse moves it, so describe only the motion and audio to add and lock the identity — "keep the face, wardrobe, and background identical." In video-to-video (V2V) you restyle or extend an existing clip; describe the new style or the continuation and, again, lock what must stay the same.

Animate this portrait: the man blinks naturally and gives a slow, warm smile, his hair shifting slightly in a gentle breeze, a subtle head turn toward the camera. Add a barely perceptible slow push-in. Keep his face, wardrobe, lighting, and background identical to the photo. Native audio: quiet room tone and a soft breath. Portrait 4:5, one short continuous shot.

Why it works: naming only small, natural motions and saying "keep identical" gives a believable live portrait without distorting the face — motion and audio only, identity locked.

Restyle this clip into a hand-painted watercolor look with soft paper texture and muted tones, keeping the exact same motion, framing, and timing as the original. Do not change the composition or the action; only change the visual style. Native audio: keep the original ambience, add a gentle acoustic guitar motif underneath. Widescreen 16:9, one continuous take.

Best for: V2V restyles where the movement is already right and you only want a new coat of paint. For the still-image side of the workflow, see how to prompt Meta Muse Image for photorealism.

Common mistakes

Most weak Muse Video output traces to one of these habits. Fix them before anything else.

  • Too many actions in one shot. "She runs, jumps, waves, and laughs while it starts to snow" gives Muse nothing clean to solve. One continuous action per shot; split the rest into Shot 2 and Shot 3.
  • Contradicting the camera. Do not ask for a "static lock-off" and a "fast tracking orbit" in the same breath. Pick one move so the frame has a single direction.
  • Adding --flags or px/fps numbers. Muse takes no --ar, no resolution codes, no "24fps" or "8s". Those tokens do nothing — state the ratio and length in words instead.
  • Over-long dialogue. A paragraph of speech will drift, because sync is still improving. Keep it to one short line in quotes per shot and let the visuals carry the rest.
  • Sound with no source. "Add dramatic sound" is vague. Name the visible object making each sound so the model can align it — "the kettle whistling", "waves breaking on the rocks".
  • A comma-tag list instead of prose. "woman, alley, neon, tracking shot, rain" throws away the relationships. Write it as a paragraph a director could follow.

Get these right and the formula does the rest. When you have the pattern down, browse the paste-ready packs: the full roundup and, when sound is the point, the Meta Muse Video prompts with audio.

Frequently Asked Questions

What is Meta Muse Video and who makes it?

Meta Muse Video is a text-to-video model from Meta Superintelligence Labs (MSL), announced on July 7, 2026 alongside Muse Image. At preview it ranked #3 for text-to-video on the LMArena human-preference leaderboard. It is built on the same pretraining base as Muse Image, generates native synchronized audio in the same pass as the picture, and is rolling out through the Meta AI app, meta.ai, Instagram Reels and Stories, and WhatsApp.

How do I write a good Meta Muse Video prompt?

Write one flowing natural-language paragraph, not a comma-tag list and not command-line flags. Name the subject and its appearance, give it one clear action, add one explicit camera move, then the setting, then lighting and mood, then an audio line, and end by stating the aspect ratio in words. Muse reasons over and self-revises the prompt before rendering, so detailed, well-structured prose pays off. Do not include numeric parameters, resolution, fps, or a seconds count.

How does native audio work in Meta Muse Video?

Sound is generated in the same pass as the picture, so dialogue, ambience, sound effects, and music all come out of one prompt. Describe the sound world in plain language and name which visible object makes each sound, for example waves breaking on rocks or a kettle whistling on the stove. Keep spoken lines short and in quotes. Meta notes audio and video sync is still improving, so do not rely on perfect lip-sync for long lines.

Does Meta Muse Video do lip-sync and dialogue?

Yes, it can generate spoken dialogue as part of the native audio, but keep lines short and put the exact words in quotes. Because sync is still improving, one brief line per shot reads far better than a paragraph of speech. Frame the speaker clearly and describe their delivery, for example she says softly, and treat perfect lip-sync as a bonus rather than a guarantee.

What is the difference between text-to-video, image-to-video, and video-to-video?

Text-to-video builds the whole clip from your written description. Image-to-video animates a still you upload — describe only the motion and audio to add and tell Muse to keep the face, wardrobe, and background identical. Video-to-video restyles or extends an existing clip. Use image-to-video when you already have the exact look locked in a frame and only want it to move, and text-to-video when you want Muse to invent the scene.

How do I set the aspect ratio and length in Meta Muse Video?

State both in words at the end of the prompt, never as numbers. Write vertical 9:16 format, a single continuous shot, or widescreen 16:9, one smooth take. Available ratios in words are square 1:1, landscape 16:9, vertical 9:16, portrait 4:5, and classic 4:3. Meta has not published exact resolution, fps, or duration figures, so describe length as one short continuous shot rather than a number of seconds.

Where can I use Meta Muse Video?

It is rolling out to creators through the Meta AI app, meta.ai on the web, Instagram Reels and Stories, and WhatsApp. Availability depends on your region and account, and features are still expanding. Because it shares a pretraining base with Muse Image, the prompting style is the same natural-language prose in both.

How is Meta Muse Video different from Meta Muse Image?

They share a pretraining base and the same prose-prompt style, so what you learn for one transfers to the other. Muse Image makes still photos and rewards camera, lens, and lighting detail. Muse Video adds motion, an explicit camera move, native synchronized audio, and multi-shot consistency, so you also describe the action, the camera direction, and the sound world, and you end with the aspect ratio and length in words.

What are the most common Meta Muse Video prompt mistakes?

Stacking too many actions into one shot, contradicting your own camera direction, adding --flags or pixel and fps numbers the model ignores, and writing dialogue that is too long to sync. Fix them by keeping one action and one camera move per shot, choosing a single coherent move, stating the ratio and length in words, and holding spoken lines to one short quoted sentence.

Advertisement