MiniMax H3 Prompt Guide: The Complete Prompting Framework for 2K Video with Native Audio (2026)
Meta Description: Master MiniMax H3 prompting with the official 6-block formula, 25 copy-ready templates, and expert techniques for text-to-video, image-to-video, and reference-to-video. Covers camera, audio, multi-shot, and troubleshooting. Updated for H3 and H3 Max.
Target Keyword: minimax h3 prompt guide Secondary Keywords: minimax h3 prompting guide, minimax h3 prompt, minimax prompt guide, hailuo 3.0 prompt, mini max h3 video prompt, h3 prompt formula
MiniMax H3 doesn’t work like other video models. Throw a pile of adjectives at Sora and you’ll get something pretty. Throw the same pile at H3 and you’ll get a confused frame that can’t decide what to animate.
H3 is a director’s model. It reads your prompt as a timed shot list: what’s in the opening frame, what moves, how the camera behaves, what it sounds like, and what the final frame looks like. Get that sequence right and H3 rewards you with 2K video, native stereo audio, and motion physics that actually make sense.
Get it wrong and you burn credits.
This guide is the definitive reference for MiniMax H3 prompting—covering every generation mode (text-to-video, image-to-video, first-and-last-frame, and reference-to-video), the official prompt structure, camera and audio vocabulary, 25 copy-ready templates, and a troubleshooting section that fixes the most common failures.
Whether you’re using H3 on fal.ai, C Dance AI, ComfyUI, or the MiniMax platform directly, this guide applies.
Table of Contents
- What Makes MiniMax H3 Different
- The 6-Block H3 Prompt Formula
- Prompting by Mode: T2V vs I2V vs FL2V vs Ref2V
- Camera Movement Reference for H3
- Audio Prompting: Dialogue, Ambience, and Music
- Multi-Shot Prompting with H3
- 25 Copy-Ready MiniMax H3 Prompt Templates
- Troubleshooting: When H3 Ignores Your Prompt
- H3 vs. Competitors: Prompt Style Differences
- MiniMax H3 Version Comparison
- FAQ: MiniMax H3 Prompting
- Conclusion
What Makes MiniMax H3 Different
MiniMax H3 (also known as Hailuo 3.0) is MiniMax’s next-generation multimodal video model. Released with open weights in July 2026, it’s designed as a general-purpose multimodal model rather than a set of separate task models.
Core Capabilities
- 2K resolution output (1440p on the short edge)
- 5 to 15 seconds per generation at 24 FPS
- Native stereo audio — dialogue, ambience, effects, and music generated alongside video
- Multi-modal input — text, images, video, and audio all in one context
- Reference-to-video — up to 9 images, 3 video clips, and 3 audio clips (12 files total)
- Legible text rendering — UI, subtitles, brand assets, and signage
- Open weights — runs locally on consumer GPUs (RTX 3060+)
What H3 Is Best At
As fal.ai’s H3 prompting guide notes: “It is strong at three things in particular: reading many references at once, rendering legible text and interfaces, and making precise localized edits to video you already have.”
This makes H3 uniquely suited for:
- Commercial production — product ads with text overlays and brand assets
- Character-driven narrative — identity-locked characters across shots
- Audio-visual content — dialogue scenes, music videos, ASMR
- Multi-reference workflows — combining face, body, environment, and style references
Where H3 Differs from Other Models
| Aspect | MiniMax H3 | Kling 3.0 | Seedance 2.0 | Sora 2 |
|---|---|---|---|---|
| Prompt style | Timed shot list | Structured 4-part | 6-block formula | Descriptive prose |
| Native audio | ✅ Stereo | ✅ Yes | ✅ Yes | ✅ Yes |
| Max resolution | 2K | 1080p | 1080p | 1080p |
| Max duration | 15s | 15s | 15s | ~25s |
| Reference inputs | 9 images + 3 video + 3 audio | Multi-image | Multi-image | Limited |
| Text rendering | ✅ Excellent | ⚠️ Moderate | ⚠️ Moderate | ⚠️ Moderate |
| Open weights | ✅ Yes | ❌ No | ❌ No | ❌ No |
The 6-Block H3 Prompt Formula
MiniMax H3 prompts work best when written as a timeline. Define the opening frame, the visible action, the camera path, the sound, and the ending state—in that order.
[Opening Frame] + [Subject Action] + [Environment Response] + [Camera Path] + [Audio] + [Final Frame]
Block 1: Opening Frame
What’s in the frame before anything moves? Define the subject, position, background, and lighting.
Text-to-video: You must build the entire opening frame from scratch. Image-to-video: The image IS the opening frame. Don’t re-describe it. Reference-to-video: Define what each reference controls.
Block 2: Subject Action
One primary action. Describe cause and effect, not just the action itself.
- ❌ “Wind blows” → vague
- ✅ “Her scarf lifts and trails behind her” → physics-accurate
Block 3: Environment Response
How does the environment react to the action? Rain on surfaces, dust rising, fabric moving, reflections shifting.
Block 4: Camera Path
One primary camera move. H3 understands push-in, pull-out, pan, tilt, track, orbit, static, crane, and more.
Block 5: Audio
H3 generates native stereo audio. Separate your audio into layers:
- Dialogue: exact lines with speaker labels
- Ambience: environment sounds
- Effects: physical sounds (footsteps, impacts, fabric)
- Music: background score (specify genre, mood, instruments)
Block 6: Final Frame
What does the last frame show? This anchors the ending state and prevents drift.
Example Prompt
Real-life studio product video. A cold citrus soda can stands
upright on a dark stone surface. A narrow stream of water
strikes the surface beside the can, splashes curl around the
base, and condensation drips down the metal. Camera orbits
slowly clockwise with the label facing the lens. Hard edge
light, restrained reflections. The shot ends with the water
settled and the can centered.
This prompt defines the opening frame (can on stone), the action (water strike and splash), the environment response (condensation, curling splash), the camera (slow clockwise orbit), the audio (implied water sounds), and the final frame (water settled, can centered).
Prompting by Mode: T2V vs I2V vs FL2V vs Ref2V
H3 supports four generation modes. Each requires different prompt behavior.
Text-to-Video (T2V)
No source media. You define everything from scratch.
Prompt must include: Full opening frame description (subject, position, background, lighting), action, camera, audio, and final state.
Common mistake: Starting with action before establishing the scene. H3 needs to know what exists before it can animate it.
A woman in a red coat stands at the edge of a foggy pier at
dawn. She takes a slow step forward, coat swaying. Camera
tracks alongside at waist height. Fog drifts across the water.
Seagulls in the distance. She stops at the end of the pier
and looks out. No music, ambient only.
Image-to-Video (I2V)
Upload an image as the first frame. The image defines the look; your prompt defines the motion.
Prompt must NOT: Re-describe what the image already shows. Prompt must: Describe what happens AFTER the first frame.
Using the uploaded image as the exact opening frame. Preserve
the woman's face, blue wool sweater, silver necklace, seat
position, and train car layout. She lifts her gaze from a
folded letter, looks toward the rain-covered window, then
folds the paper along its existing crease. Camera slowly
tracks right as her reflection crosses the glass. Rain taps
the window, overtaking the low rhythm of train wheels. No
background music.
First-and-Last-Frame (FL2V)
Upload both a starting and ending image. H3 generates the motion between them.
Prompt must: Describe the physical transition between frames—how posture, objects, camera, and lighting change from start to end.
Common mistake: Describing two static images. H3 needs the path between them.
The subject rises from the chair, takes three steps toward
the window, and turns to face camera. Camera pulls back
from medium close-up to medium wide. Lighting shifts from
warm interior to cool window light as she moves. Ambient
room tone transitions to street sounds from the window.
Reference-to-Video (Ref2V)
Upload multiple references (images, video, audio) and assign each a specific role.
Prompt must: Explicitly state what each reference controls.
@Image1 defines the character's face and hair.
@Image2 defines the clothing and color palette.
@Video1 provides the body motion and camera rhythm.
@Audio1 provides the voice timing.
Use @Image1 person as the sole performer. Preserve her
facial proportions, short black hair, and red jacket from
@Image1. Follow the body action and shot timing from @Video1,
but place the action in the subway platform shown in @Image2.
Keep the camera path and performer pose sequence from @Video1.
Use @Audio1 for timing; do not add background music.
Key rule: Conflicting references cause drift. If two images show different clothing or face shapes, H3 must guess which to keep. Remove the weaker reference before adding more prompt text.
Camera Movement Reference for H3
H3 understands specific camera verbs. Use them precisely.
| Movement | Prompt Phrase | Best For |
|---|---|---|
| Push in | “Camera slowly pushes in toward [subject]” | Reveals, tension, emphasis |
| Pull out | “Camera slowly pulls back to reveal [scene]” | Establishing context |
| Pan left/right | “Camera pans [direction] across the scene” | Horizontal reveals |
| Tilt up/down | “Camera tilts [direction]” | Vertical reveals |
| Track | “Camera tracks alongside [subject]” | Walking, movement sequences |
| Orbit | “Camera orbits around [subject]” | 360-degree views |
| Static | “Camera remains static and locked off” | Product shots, talking heads |
| Crane up/down | “Camera cranes [direction]” | Dramatic vertical shifts |
| Handheld | “Subtle handheld movement” | Documentary, authentic feel |
| Whip | “Quick whip pan to [subject]” | Transitions, energy |
| Zoom | “Slow zoom in on [detail]” | Detail emphasis |
H3-specific tips:
- Use one primary camera move per shot. Stacking moves (“orbit + zoom + whip”) destabilizes output.
- Write camera as a scene instruction, not a tag list: “Camera tracks alongside the cyclist” beats “tracking shot, dynamic camera.”
- For reveals, describe the starting view, the movement, and what’s revealed: “Close-up starts on the baker’s hands. Lens slowly pulls back to reveal the entire workbench.”
Audio Prompting: Dialogue, Ambience, and Music
H3 generates native stereo audio alongside video. Separate your audio into three distinct layers:
Layer 1: Dialogue
Put exact lines in quotes. Label the speaker. Specify language and tone.
[Speaker 1: woman, warm voice, English]: "I've been waiting
for this moment."
[Speaker 2: man, quiet tone, English]: "So have I."
Tips:
- Keep lines short. Time them aloud—if a sentence takes 5 seconds, it won’t fit naturally in a 6-second clip with action.
- Specify “their lips only move during their own lines” to prevent cross-sync.
- For voiceover, state it’s off-screen and the on-screen subject’s lips stay closed.
Layer 2: Ambience & Effects
Name the source, timing, and volume of physical sounds.
Sound: steady wheel noise, light rain on glass, gentle
ticket-punching sound. Footsteps on gravel. Door creaking
open. Fabric rustling in wind.
Layer 3: Background Music
Describe genre, mood, instruments, and when it enters/exits.
Background music: sparse piano notes at a slow tempo, fading
in after the character speaks, building gently toward the end.
Or silence:
No background music. Ambient sound only.
Multi-Shot Prompting with H3
H3 supports multi-shot generation within a single prompt. Use [Shot N] labels and timestamp-based cuts.
Format
[Shot 1] Opening shot description. Duration and action.
[Shot 2] At 00:03.000, cut to next shot. Description.
[Shot 3] At 00:07.000, cut to final shot. Description.
Example: 10-Second Product Ad (3 Shots)
[Shot 1] Real-life ad, extreme close-up of a runner tying
shoelaces on a wet track at pre-dawn. Camera low and static.
Fabric tightens, lace pulls taut. 3 seconds.
[Shot 2] At 00:03.000, cut to side tracking shot as the runner
accelerates through the first curve. Each footstep splashes
a small spray from the track. Breathing and foot impact
overtake distant city ambience. 4 seconds.
[Shot 3] At 00:07.000, cut to the shoe landing beside the
finish line, then slow push-in as the runner stops. Brief
bass pulse ends on the final footstep. 3 seconds.
Multi-Shot Best Practices
- Allocate time deliberately — 3s for setup, 4s for action, 3s for ending
- Keep each shot simple — one action, one camera move per shot
- Use timestamps for cuts, not for continuous shots
- Maintain identity — reuse exact character descriptions across shots
- Audio continuity — describe how sound transitions between shots
- Don’t timestamp a single continuous shot — use timestamps only for cuts between shots
25 Copy-Ready MiniMax H3 Prompt Templates
Product & Commercial (5 Templates)
1. Product Hero with Liquid
Real-life studio product video. A [product] stands on a
[ surface]. A narrow stream of [liquid] strikes the surface
beside it, splashes curling around the base. Camera orbits
slowly clockwise with the label facing the lens. Hard edge
light, restrained reflections. The shot ends with the liquid
settled and the product centered.
2. Product Reveal
[Shot 1] Close-up of a branded box on a dark surface.
Hands slowly open the lid, revealing [product] inside.
Soft overhead light fills the box. 4 seconds.
[Shot 2] Camera slowly pulls back to reveal the full product.
Hands lift it out and turn it to face camera. Clean studio
lighting. 6 seconds.
3. Text Overlay Commercial
Real-life product commercial. [Product] on [surface].
Animated text "[BRAND TEXT]" slides in from below and locks
into position with a subtle bounce. Camera static. Clean
modern design. Motion graphics style. No text distortion.
4. Watch Macro Detail
Extreme macro of a luxury watch face on dark marble. The
second hand sweeps smoothly. Hard pin-spot from above
catches the polished bezel. Camera slowly pushes in toward
the dial center. Tick-tock sound, no music. 8 seconds.
5. Jewelry Rotation
A diamond ring on a black velvet pedestal rotates slowly
on a turntable. Single hard spotlight from upper right
catches each facet. Camera static close-up. Deep black
and cool white sparkle palette. 7 seconds.
Character & Dialogue (5 Templates)
6. Character Walking (Physics-Accurate)
Low-angle tracking shot at street level. A [character
description] walks through [location]. Steady pace. Each
step lands heel-first, rolling forward with visible weight
transfer. [Surface] reflects [lighting]. 35mm film, shallow
depth of field. Ambient [sound].
7. Two-Person Dialogue
[Shot 1] Medium close-up of [Character A] facing camera.
[Character A: description, voice]: "[dialogue]." Natural
[lighting]. 4 seconds.
[Shot 2] Cut to [Character B] responding.
[Character B: description, voice]: "[dialogue]." Same
lighting setup. 4 seconds.
8. Emotional Close-Up
Extreme close-up of [character]'s face. [Emotion] crosses
their features slowly—eyes soften, shoulders drop slightly.
Single hard key from camera-[direction]. Shallow depth of
field. Intimate, vulnerable. No face morphing. 6 seconds.
9. Character Interaction with Object
[Character] reaches out and picks up [object] from [surface].
Fingers wrap around it with natural grip. Camera slowly
pushes in. [Lighting]. Cinematic [style]. No hand deformation,
no object distortion. 6 seconds.
10. Character with Wind
[Character] stands at [location]. Gentle wind from camera-
[right] lifts and flows their [hair/clothing] naturally.
Camera static medium shot. Golden hour backlight. Cinematic.
No hair merging, no face distortion. 6 seconds.
UGC & Social Media (5 Templates)
11. Creator Talking Head
Vertical phone video. Creator faces camera in a bright room.
"[Hook line]" in an enthusiastic, genuine tone. Natural
window light. Slight handheld movement. Social-first pacing.
No face distortion. 8 seconds.
12. Unboxing
Overhead shot. Hands open a branded box on a table, lift
out [product], and turn it to show the camera. Natural
indoor lighting. UGC style. "Look what just arrived" in an
excited whisper. No hand clipping. 8 seconds.
13. Before/After Reveal
[Shot 1] [Before state]. Camera static. Clean background.
3 seconds.
[Shot 2] Quick cut to [after state]. Same framing. Bright
studio lighting. Satisfying transformation moment. 4 seconds.
14. Product Demo
Creator holds [product] to camera. Demonstrates [feature]
with natural hand motion. "[Product demo line]" in a
confident, helpful tone. Warm studio lighting. Social
pacing. No hand deformation. 10 seconds.
15. Flat Lay Animation
Overhead shot of a styled flat lay arrangement. One element
slowly slides into frame from the edge and settles into
position. Camera static. Soft even daylight. Editorial
style. No element overlap. 6 seconds.
Cinematic & Nature (5 Templates)
16. Establishing Shot
Wide establishing shot of [location] at [time of day].
[Atmospheric element] drifts across the scene. Camera
slowly cranes up revealing the full scale. Cinematic [color]
palette, volumetric light. Ambient [sound], no dialogue.
10 seconds.
17. Storm Sequence
[Shot 1] Wide shot of [location] as storm approaches.
Darkening sky, wind picks up. 4 seconds.
[Shot 2] Rain begins, lightning flash illuminates the scene.
Camera static. Deep thunder. 4 seconds.
[Shot 3] Downpour. Camera slowly pushes in toward [subject].
Rain on surfaces. 4 seconds.
18. Golden Hour Landscape
Wide shot of [landscape] at golden hour. Long shadows stretch
across the terrain. Warm low sun. Camera slow crane up
revealing the full vista. Gentle wind moves vegetation.
Ambient natural sound. No music. 10 seconds.
19. Foggy Forest Walk
Medium wide shot moving slowly through a dense forest at
dawn. Thick fog drifts between trees. Volumetric light rays
through the canopy. Camera tracks forward at walking pace.
Muted green and grey palette. Soft ambient forest sounds.
10 seconds.
20. Underwater Scene
Wide shot underwater. [Subject] moves through the water
with natural buoyancy. Dappled light from the surface
creates moving caustics. Camera slowly tracks alongside.
Dreamy pace. Aquatic ambient sound. 8 seconds.
Food & Beverage (5 Templates)
21. Coffee Pour
Close-up of a ceramic cup on a wooden surface. Coffee pours
from above in a smooth stream, filling the cup. Steam rises
slowly. Camera slowly pushes in. Warm café lighting, soft
bokeh background. Café ambience. 6 seconds.
22. Sizzling Dish
Overhead close-up of [dish] on a hot plate. [Ingredient]
sizzles and pops. Steam rises. Camera slowly pushes in.
Warm kitchen lighting. Sizzling sound, no music. 6 seconds.
23. Cocktail Assembly
Close-up of a glass on a bar surface. [Liquid] pours over
ice in slow motion. Ice cracks and shifts. Camera slowly
orbits. Dramatic bar lighting. Liquid sound, clinking ice.
No liquid clipping. 7 seconds.
24. Ingredient Prep
Close-up of hands slicing [ingredient] on a wooden cutting
board. Each cut is precise and rhythmic. Overhead angle.
Bright natural kitchen lighting. Knife-on-wood sound.
No hand deformation. 6 seconds.
25. Chocolate Break
Close-up of a chocolate bar. A hand slowly breaks off a
piece, revealing the cross-section. The piece separates
cleanly with a satisfying snap. Camera static macro. Warm
studio lighting. Snap sound, no music. 5 seconds.
For more tested H3 prompts with their generated outputs, browse the VideosPrompt community which features prompts organized by model and category.
Troubleshooting: When H3 Ignores Your Prompt
When output fails, change the minimum instruction that explains the failure. Don’t rewrite the entire prompt.
| Failure | Likely Cause | First Fix |
|---|---|---|
| Random cuts | Multiple camera moves or scene changes competing | Request one continuous shot, one camera path |
| Wrong person speaking | Speaker identification inconsistent | Assign stable labels to each speaker, shorten dialogue |
| Dialogue feels rushed | Lines too long for the duration | Time lines aloud, or increase clip duration |
| Product deforms | Too many actions, camera, and style elements | Lock product orientation, remove secondary motion |
| Reference identity drift | Conflicting references or unclear character | Keep the clearest identity image, state remaining characters explicitly |
| First and last frame don’t connect | Missing physical transition | Describe the intermediate poses, objects, and camera changes |
| Audio feels unrelated | Audio described as mood only | Name the source, timing, and volume changes |
| Subject ignores camera | Camera instruction buried in description | Move camera instruction to its own sentence, early in the prompt |
| Style inconsistency | Multiple style descriptors fighting | Use one style label: “cinematic” or “documentary,” not both |
Key principle: If the camera is wrong but the subject is correct, keep the subject language and modify only the camera sentence. Each retry should change one variable.
H3 vs. Competitors: Prompt Style Differences
The same concept needs different prompt language for each model:
Same Scene, Four Dialects
Concept: A woman walks through a rain-soaked alley at night.
MiniMax H3:
Real-life cinematic shot. A woman in a dark coat walks
through a rain-soaked alley at night. Steady pace, each
step splashes in puddles. Camera tracks alongside at waist
height. Neon reflections on wet pavement. Ambient rain,
distant traffic. She stops at the end of the alley and
looks up. No music.
Kling 3.0:
Medium tracking shot. Woman in dark coat walks through
rain-soaked alley at night. Steady pace, puddle splashes.
Neon reflections on wet pavement. 35mm film, shallow DOF.
Ambient rain, distant traffic.
Seedance 2.0:
Woman in dark coat walking through rain-soaked neon alley
at night. Side tracking shot, steady pace. Wet pavement
reflections. Cinematic, moody. Rain ambience, no music.
Sora 2:
A woman in a dark overcoat walks steadily through a narrow
rain-soaked alley at night. Neon signs reflect off the wet
pavement in pools of color. The camera follows her from the
side at waist height. Rain falls gently. Cinematic, moody
atmosphere.
H3 prompts read like timed shot lists. Kling reads like structured controls. Seedance reads like scene briefs. Sora reads like descriptive prose.
MiniMax H3 Version Comparison
| Version | Resolution | Duration | Audio | Key Feature | Best For |
|---|---|---|---|---|---|
| H3 (Base) | Up to 2K | 5-15s | Stereo | Open weights, multi-modal | Local deployment, experimentation |
| H3 Max (fal) | Up to 2K | 5-15s | Enhanced stereo | Post-trained, Director endpoint | Production quality, commercial |
| H3 Turbo | 768p | 5-10s | Stereo | Fast generation | Drafts, batch testing |
When to use each:
- H3 Base: Open-source workflows, ComfyUI, local GPU deployment
- H3 Max: Production commercials, brand films, hero content
- H3 Turbo: Rapid iteration, A/B testing, draft previews
FAQ: MiniMax H3 Prompting
What is the MiniMax H3 prompt formula?
The 6-block formula: [Opening Frame] + [Subject Action] + [Environment Response] + [Camera Path] + [Audio] + [Final Frame]. This sequence gives H3 a timeline to follow rather than a pile of adjectives to interpret.
How long should a MiniMax H3 prompt be?
H3 accepts up to 7,000 characters, but most effective prompts are 3-6 sentences. Longer isn’t better—clearer is better. Each sentence should specify one dimension: opening frame, action, camera, audio, or final state.
How do I write dialogue for MiniMax H3?
Put exact lines in quotes. Label each speaker with description and voice tone. Keep lines short—time them aloud. Specify language. For multi-speaker scenes, add “their lips only move during their own lines” to prevent cross-sync issues.
What’s the difference between H3 text-to-video and image-to-video?
Text-to-video requires you to define the entire opening frame from scratch. Image-to-video uses your uploaded image as the opening frame—your prompt should only describe what happens AFTER the image, not re-describe what’s already visible.
How many reference files can H3 accept?
Up to 12 files total: 9 images, 3 video clips (2-15 seconds each), and 3 audio clips. Each reference should have an explicit role—don’t upload conflicting references that ask the model to guess which to follow.
Can H3 render text and UI elements?
Yes. H3 is notably strong at rendering legible text, subtitles, brand logos, and UI elements. Specify the exact text you want rendered and its position on screen.
Why does H3 ignore parts of my prompt?
Common causes: too many stacked actions, conflicting camera moves, ambiguous reference characters, or dialogue that’s too long for the duration. Remove secondary instructions and test one change at a time.
Where can I use MiniMax H3?
H3 is available on fal.ai (API, pay-per-use), C Dance AI (web interface), ComfyUI (local deployment with open weights), and the MiniMax platform. The VideosPrompt community also features H3 prompts you can study and adapt.
Conclusion
MiniMax H3 is a director’s model. It doesn’t want adjectives—it wants a shot list.
The framework is simple:
- Define the opening frame — what exists before motion starts
- Describe one action — with cause and effect, not just the verb
- Set the camera — one move, one path, one instruction
- Direct the sound — dialogue, ambience, effects, music as separate layers
- Anchor the final frame — where things end up
- Protect fragile elements — explicit constraints for text, hands, faces
Start with the templates above. Swap in your subjects. Generate, compare, refine.
Every clip builds your intuition for how H3 interprets timed instructions. After 50 generations, you won’t need templates—you’ll look at a scene and see the shot list underneath.
Your stories deserve 2K resolution and stereo sound. Now you know how to direct them.
Last updated: September 2026
Ready to explore more? Browse the VideosPrompt community for thousands of curated AI video prompts across every model—including dedicated MiniMax H3 collections. Or try SocialToPrompt to reverse-engineer any video into a reusable prompt.
Share Article