Hunyuan Image 3.0 is Tencent's open-source image model — a native multimodal, unified autoregressive model with 80 billion parameters (about 13B activated) across a 64-expert mixture, and the largest open-source image model at launch. This is the quick-reference version: the six-part prompt formula, then copy-paste tables for aspect ratios, lighting, lens, style, and text, and a few complete example prompts. For the full pack, keep the best Hunyuan Image prompts open alongside this sheet.

Two facts shape everything below. Hunyuan has no inline parameters — there is no --ar, --v, or --style, so ratio and resolution are settings you pick in the interface or API. And its world-knowledge reasoning plus best-in-class bilingual text mean you can lean on plain natural language: describe the scene, quote any words, and it fills accurate detail. For the lens and lighting language behind photo work, see the photorealism guide.

Advertisement

The prompt formula

Write one natural sentence in this order and Hunyuan reads it cleanly. Any words that should appear inside the image go in double quotes.

PartWhat to writeExample fragment
1. SubjectWho or what, with a defining detaila potter in her sixties with clay-dusted hands
2. Action / poseWhat they are doing or how they sitlooking up from her wheel
3. SettingWhere, and the surrounding detailin a warm sunlit studio with shelves of bowls
4. LightingNamed light setup and time of daysoft directional window light with a warm bounce
5. Camera / lensFocal length and aperture, if photorealshot on a 35mm lens, shallow depth of field
6. StyleMedium or aesthetic to commit tonatural documentary photography, honest color
+ TextExact words in "quotes", short and namedheadline "JADE & STEAM" in a bold serif

Aspect ratios & resolution

Ratio and size are UI or API settings, not --flags. The maximum output is 2048x2048 (2K) — Hunyuan does not natively produce 4K; upscale afterward if you need more.

Aspect ratioOrientationBest for
1:1SquareLogos, app icons, product hero, social posts
4:3LandscapeInteriors, infographics, classic photo frame
3:4PortraitInk-wash scenes, book pages, standing subjects
3:2LandscapeLandscapes, editorial photography, watch macro
2:3PortraitPosters, magazine covers, fashion editorial
16:9Wide landscapeBanners, key art, cinematic scenes, wallpapers
9:16Tall portraitPhone wallpapers, stories, vertical video frames

Pick the ratio and a size up to 2K in the tool; you can also name the ratio in plain English inside the prompt as a hint.

Lighting modifiers

Name the light first — it sets mood faster than anything else. Drop one of these phrases into part four of the formula.

ModifierEffect
Golden hourWarm, low, raking sun with long soft shadows
Blue hourCool dusk glow, deep blue sky, lit foreground
Soft window lightGentle directional daylight, natural falloff
Three-point lightingClean, controlled studio look with key, fill, rim
High-keyBright, even, low-contrast, minimal shadow
Low-key / chiaroscuroDeep blacks, one bright highlight, dramatic
Overcast / diffusedEven soft light, no harsh shadows, true color
Rim lightBright edge that separates subject from background
BacklightLight behind subject, glowing hair, hazy bokeh
Neon / practicalColored in-scene sources, moody urban reflections

Camera & lens modifiers

Focal length and aperture control framing and focus. Use these only for photoreal work; drop them for illustration or vector styles.

Lens / settingLook
24mm wideSweeping landscapes, interiors, dramatic scale
35mmNatural documentary and environmental framing
50mmNeutral, true-to-eye perspective, everyday shots
85mmFlattering portraits, gentle background compression
100mm macroExtreme close-ups of product and material texture
f/1.4 (shallow DOF)Blurred background, strong subject separation
f/8 (deep DOF)Sharp front to back, best for landscapes
Tilt-shiftStraight architecture lines or a miniature effect
Advertisement

Style & medium modifiers

Name the medium and Hunyuan commits to it. Its cultural training makes ink-wash and traditional Asian styles especially convincing.

StyleCue words
Photorealisticphotorealistic, natural color, fine skin/texture detail
Cinematiccinematic color grade, film grain, dramatic lighting
Anime / cel-shadedcel-shaded anime, clean linework, vibrant flat color
Ink-wash / shui-moChinese ink-wash, loose brushwork, rice-paper texture
Watercolorsoft blooms, visible paper texture, muted palette
Oil paintingvisible brushstrokes, rich impasto, painterly light
3D low-poly / isometricisometric render, soft global illumination, pastel
Flat vectorflat-shaded vector, bold outlines, no gradients/shadow
Concept artdigital concept art, dramatic composition, moody color

Text & typography rules

In-image text is a signature strength — legible English and Chinese in the same image. Follow these rules for clean, correctly spelled type.

RuleWhy
Put exact words in "double quotes"Tells Hunyuan which strings to render literally
Keep each string 2–10 wordsShort copy stays legible and correctly spelled
Name the font style"bold condensed sans-serif" guides the letterforms
Describe placement"top-center", "along the bottom banner" fixes layout
Bilingual English + Chinese supportedQuote a Latin headline and a Chinese subtitle together
Isolate a misspelled word and regenerateShorten and re-quote rather than piling on more copy

For a full set of typography-first prompts, see the Hunyuan prompt templates.

Settings-line convention & example prompts

Because Hunyuan has no inline --parameters, these packs add a short Settings reminder line under each prompt — for example Settings: Hunyuan Image 3.0 · 16:9 · 2K — telling you which model, ratio, and size to select in the tool. Here are four complete prompts using the formula above.

A fisherman in his fifties mending a net on a weathered wooden dock at golden hour, warm low sunlight raking across his face and the coiled rope, calm harbor and moored boats softly blurred behind him, natural documentary photography with honest color and fine skin texture, shallow depth of field, shot on an 85mm lens at f1.8.

Settings: Hunyuan Image 3.0 · 3:2 · 2K

A clean festival poster with the English headline "SPRING MARKET" in a bold rounded sans-serif across the top and the Chinese subtitle "春季市集" directly below in a matching weight, plus a footer line "SAT MAR 21 - 10AM". Flat illustration of paper lanterns and blossom branches on a warm cream background, generous spacing, all text sharp and correctly rendered in both languages.

Settings: Hunyuan Image 3.0 · 2:3 · 2K

A traditional Chinese ink-wash painting of a lone fisherman poling a small boat across a misty lake at dawn, distant peaks fading into pale washes of grey and black, a single reed in the foreground, expressive loose brushwork with visible rice-paper texture and generous negative space, elegant and poetic, minimal palette with one faint touch of warm ochre.

Settings: Hunyuan Image 3.0 · 3:4 · 2K

Replace the plain grey background behind the product with a softly blurred marble kitchen counter in warm morning light, keep the ceramic mug, its handle, glaze, and the steam rising from it exactly the same, match the new background's warm color temperature and add a natural contact shadow beneath the mug so it sits realistically, photorealistic result with clean edges.

Settings: Hunyuan Image 3.0-Instruct · edit · 1 reference image · keep subject identical

Frequently Asked Questions

Is there an --ar flag in Hunyuan Image?

No. Hunyuan Image 3.0 has no inline parameters like --ar, --v, or --style. Aspect ratio and resolution are interface or API settings you pick in the tool, not text you type in the prompt. Because there are no flags, these packs add a short Settings reminder line under each prompt — for example "Settings: Hunyuan Image 3.0 · 16:9 · 2K" — so you know which ratio and size to select. You can also mention the ratio in plain English inside the prompt.

What is the maximum resolution of Hunyuan Image 3.0?

The maximum output resolution is 2048x2048, which is 2K. Hunyuan Image 3.0 does not natively produce 4K. Choose an aspect ratio (1:1, 4:3, 3:4, 3:2, 2:3, 16:9, or 9:16) and a size up to that 2K ceiling in the interface or API. If you need larger files, generate at 2K and upscale afterward with a separate tool.

How do I get clean, correctly spelled text in an image?

Put the exact words in double quotes, keep each string short (roughly two to ten words), name the font style, and say where the text sits on the canvas. Hunyuan Image 3.0 renders both English and Chinese characters accurately, so bilingual posters and packaging are a real strength. If a word comes out wrong, isolate it as its own quoted element, shorten the copy, and regenerate rather than adding more text.

Does Hunyuan Image handle Chinese text?

Yes. Best-in-class bilingual English and Chinese in-image text is one of Hunyuan Image 3.0's headline strengths. You can quote a Latin headline and a Chinese subtitle in the same prompt and it renders both legibly, which makes it a strong pick for bilingual flyers, packaging, and posters. It also brings strong Asian cultural authenticity to the imagery around the type.

Should I write short or long prompts?

Both work. Because Hunyuan Image 3.0 is a native multimodal model with world-knowledge reasoning, a short prompt — a named landmark, dish, or era — still resolves into a detailed, plausible scene. It also understands very long prompts (well over a thousand characters), so when you need control, write a full paragraph covering subject, setting, lighting, lens, and style, and it will honor the specifics instead of dropping them.

Is Hunyuan Image 3.0 free and open-source?

Yes. Tencent released Hunyuan Image 3.0 as open-source under an Apache 2.0 license, so you can download the weights from Hugging Face or GitHub (Tencent-Hunyuan/HunyuanImage-3.0) and run them yourself for free, including commercial use, if you have the hardware. Hosted options — the official portal at hunyuan.tencent.com, Tencent Cloud, and providers like WaveSpeed AI — may be free within limits or metered per image.

What is the difference between Hunyuan Image 3.0 and 3.0-Instruct?

The base Hunyuan Image 3.0 model generates images from text. The Instruct variant is instruction-tuned and adds reasoning-based editing on top: it can edit an existing image from a plain-language instruction, apply style transfers, and fuse several reference images into one composite. For edits, feed the image plus a clear instruction and say what to keep identical; for fusion, describe which element comes from which reference.

Advertisement