VideosPrompt VideosPrompt

MiniMax H3 Prompt Guide: The Complete Prompting Framework for 2K Video with Native Audio (2026)

Author: VideosPrompt Date: 2026-09-14 07:43:11
MiniMax H3 Prompt Guide: The Complete Prompting Framework for 2K Video with Native Audio (2026)

Meta Description: Master MiniMax H3 prompting with the official 6-block formula, 25 copy-ready templates, and expert techniques for text-to-video, image-to-video, and reference-to-video. Covers camera, audio, multi-shot, and troubleshooting. Updated for H3 and H3 Max.

Target Keyword: minimax h3 prompt guide Secondary Keywords: minimax h3 prompting guide, minimax h3 prompt, minimax prompt guide, hailuo 3.0 prompt, mini max h3 video prompt, h3 prompt formula


MiniMax H3 doesn’t work like other video models. Throw a pile of adjectives at Sora and you’ll get something pretty. Throw the same pile at H3 and you’ll get a confused frame that can’t decide what to animate.

H3 is a director’s model. It reads your prompt as a timed shot list: what’s in the opening frame, what moves, how the camera behaves, what it sounds like, and what the final frame looks like. Get that sequence right and H3 rewards you with 2K video, native stereo audio, and motion physics that actually make sense.

Get it wrong and you burn credits.

This guide is the definitive reference for MiniMax H3 prompting—covering every generation mode (text-to-video, image-to-video, first-and-last-frame, and reference-to-video), the official prompt structure, camera and audio vocabulary, 25 copy-ready templates, and a troubleshooting section that fixes the most common failures.

Whether you’re using H3 on fal.ai, C Dance AI, ComfyUI, or the MiniMax platform directly, this guide applies.


Table of Contents

  1. What Makes MiniMax H3 Different
  2. The 6-Block H3 Prompt Formula
  3. Prompting by Mode: T2V vs I2V vs FL2V vs Ref2V
  4. Camera Movement Reference for H3
  5. Audio Prompting: Dialogue, Ambience, and Music
  6. Multi-Shot Prompting with H3
  7. 25 Copy-Ready MiniMax H3 Prompt Templates
  8. Troubleshooting: When H3 Ignores Your Prompt
  9. H3 vs. Competitors: Prompt Style Differences
  10. MiniMax H3 Version Comparison
  11. FAQ: MiniMax H3 Prompting
  12. Conclusion

What Makes MiniMax H3 Different

MiniMax H3 (also known as Hailuo 3.0) is MiniMax’s next-generation multimodal video model. Released with open weights in July 2026, it’s designed as a general-purpose multimodal model rather than a set of separate task models.

Core Capabilities

  • 2K resolution output (1440p on the short edge)
  • 5 to 15 seconds per generation at 24 FPS
  • Native stereo audio — dialogue, ambience, effects, and music generated alongside video
  • Multi-modal input — text, images, video, and audio all in one context
  • Reference-to-video — up to 9 images, 3 video clips, and 3 audio clips (12 files total)
  • Legible text rendering — UI, subtitles, brand assets, and signage
  • Open weights — runs locally on consumer GPUs (RTX 3060+)

What H3 Is Best At

As fal.ai’s H3 prompting guide notes: “It is strong at three things in particular: reading many references at once, rendering legible text and interfaces, and making precise localized edits to video you already have.”

This makes H3 uniquely suited for:

  • Commercial production — product ads with text overlays and brand assets
  • Character-driven narrative — identity-locked characters across shots
  • Audio-visual content — dialogue scenes, music videos, ASMR
  • Multi-reference workflows — combining face, body, environment, and style references

Where H3 Differs from Other Models

Aspect MiniMax H3 Kling 3.0 Seedance 2.0 Sora 2
Prompt style Timed shot list Structured 4-part 6-block formula Descriptive prose
Native audio ✅ Stereo ✅ Yes ✅ Yes ✅ Yes
Max resolution 2K 1080p 1080p 1080p
Max duration 15s 15s 15s ~25s
Reference inputs 9 images + 3 video + 3 audio Multi-image Multi-image Limited
Text rendering ✅ Excellent ⚠️ Moderate ⚠️ Moderate ⚠️ Moderate
Open weights ✅ Yes ❌ No ❌ No ❌ No

The 6-Block H3 Prompt Formula

MiniMax H3 prompts work best when written as a timeline. Define the opening frame, the visible action, the camera path, the sound, and the ending state—in that order.

[Opening Frame] + [Subject Action] + [Environment Response] + [Camera Path] + [Audio] + [Final Frame]

Block 1: Opening Frame

What’s in the frame before anything moves? Define the subject, position, background, and lighting.

Text-to-video: You must build the entire opening frame from scratch. Image-to-video: The image IS the opening frame. Don’t re-describe it. Reference-to-video: Define what each reference controls.

Block 2: Subject Action

One primary action. Describe cause and effect, not just the action itself.

  • ❌ “Wind blows” → vague
  • ✅ “Her scarf lifts and trails behind her” → physics-accurate

Block 3: Environment Response

How does the environment react to the action? Rain on surfaces, dust rising, fabric moving, reflections shifting.

Block 4: Camera Path

One primary camera move. H3 understands push-in, pull-out, pan, tilt, track, orbit, static, crane, and more.

Block 5: Audio

H3 generates native stereo audio. Separate your audio into layers:

  • Dialogue: exact lines with speaker labels
  • Ambience: environment sounds
  • Effects: physical sounds (footsteps, impacts, fabric)
  • Music: background score (specify genre, mood, instruments)

Block 6: Final Frame

What does the last frame show? This anchors the ending state and prevents drift.

Example Prompt

Real-life studio product video. A cold citrus soda can stands
upright on a dark stone surface. A narrow stream of water
strikes the surface beside the can, splashes curl around the
base, and condensation drips down the metal. Camera orbits
slowly clockwise with the label facing the lens. Hard edge
light, restrained reflections. The shot ends with the water
settled and the can centered.

This prompt defines the opening frame (can on stone), the action (water strike and splash), the environment response (condensation, curling splash), the camera (slow clockwise orbit), the audio (implied water sounds), and the final frame (water settled, can centered).


Prompting by Mode: T2V vs I2V vs FL2V vs Ref2V

H3 supports four generation modes. Each requires different prompt behavior.

Text-to-Video (T2V)

No source media. You define everything from scratch.

Prompt must include: Full opening frame description (subject, position, background, lighting), action, camera, audio, and final state.

Common mistake: Starting with action before establishing the scene. H3 needs to know what exists before it can animate it.

A woman in a red coat stands at the edge of a foggy pier at
dawn. She takes a slow step forward, coat swaying. Camera
tracks alongside at waist height. Fog drifts across the water.
Seagulls in the distance. She stops at the end of the pier
and looks out. No music, ambient only.

Image-to-Video (I2V)

Upload an image as the first frame. The image defines the look; your prompt defines the motion.

Prompt must NOT: Re-describe what the image already shows. Prompt must: Describe what happens AFTER the first frame.

Using the uploaded image as the exact opening frame. Preserve
the woman's face, blue wool sweater, silver necklace, seat
position, and train car layout. She lifts her gaze from a
folded letter, looks toward the rain-covered window, then
folds the paper along its existing crease. Camera slowly
tracks right as her reflection crosses the glass. Rain taps
the window, overtaking the low rhythm of train wheels. No
background music.

First-and-Last-Frame (FL2V)

Upload both a starting and ending image. H3 generates the motion between them.

Prompt must: Describe the physical transition between frames—how posture, objects, camera, and lighting change from start to end.

Common mistake: Describing two static images. H3 needs the path between them.

The subject rises from the chair, takes three steps toward
the window, and turns to face camera. Camera pulls back
from medium close-up to medium wide. Lighting shifts from
warm interior to cool window light as she moves. Ambient
room tone transitions to street sounds from the window.

Reference-to-Video (Ref2V)

Upload multiple references (images, video, audio) and assign each a specific role.

Prompt must: Explicitly state what each reference controls.

@Image1 defines the character's face and hair.
@Image2 defines the clothing and color palette.
@Video1 provides the body motion and camera rhythm.
@Audio1 provides the voice timing.

Use @Image1 person as the sole performer. Preserve her
facial proportions, short black hair, and red jacket from
@Image1. Follow the body action and shot timing from @Video1,
but place the action in the subway platform shown in @Image2.
Keep the camera path and performer pose sequence from @Video1.
Use @Audio1 for timing; do not add background music.

Key rule: Conflicting references cause drift. If two images show different clothing or face shapes, H3 must guess which to keep. Remove the weaker reference before adding more prompt text.


Camera Movement Reference for H3

H3 understands specific camera verbs. Use them precisely.

Movement Prompt Phrase Best For
Push in “Camera slowly pushes in toward [subject]” Reveals, tension, emphasis
Pull out “Camera slowly pulls back to reveal [scene]” Establishing context
Pan left/right “Camera pans [direction] across the scene” Horizontal reveals
Tilt up/down “Camera tilts [direction]” Vertical reveals
Track “Camera tracks alongside [subject]” Walking, movement sequences
Orbit “Camera orbits around [subject]” 360-degree views
Static “Camera remains static and locked off” Product shots, talking heads
Crane up/down “Camera cranes [direction]” Dramatic vertical shifts
Handheld “Subtle handheld movement” Documentary, authentic feel
Whip “Quick whip pan to [subject]” Transitions, energy
Zoom “Slow zoom in on [detail]” Detail emphasis

H3-specific tips:

  • Use one primary camera move per shot. Stacking moves (“orbit + zoom + whip”) destabilizes output.
  • Write camera as a scene instruction, not a tag list: “Camera tracks alongside the cyclist” beats “tracking shot, dynamic camera.”
  • For reveals, describe the starting view, the movement, and what’s revealed: “Close-up starts on the baker’s hands. Lens slowly pulls back to reveal the entire workbench.”

Audio Prompting: Dialogue, Ambience, and Music

H3 generates native stereo audio alongside video. Separate your audio into three distinct layers:

Layer 1: Dialogue

Put exact lines in quotes. Label the speaker. Specify language and tone.

[Speaker 1: woman, warm voice, English]: "I've been waiting
for this moment."

[Speaker 2: man, quiet tone, English]: "So have I."

Tips:

  • Keep lines short. Time them aloud—if a sentence takes 5 seconds, it won’t fit naturally in a 6-second clip with action.
  • Specify “their lips only move during their own lines” to prevent cross-sync.
  • For voiceover, state it’s off-screen and the on-screen subject’s lips stay closed.

Layer 2: Ambience & Effects

Name the source, timing, and volume of physical sounds.

Sound: steady wheel noise, light rain on glass, gentle
ticket-punching sound. Footsteps on gravel. Door creaking
open. Fabric rustling in wind.

Layer 3: Background Music

Describe genre, mood, instruments, and when it enters/exits.

Background music: sparse piano notes at a slow tempo, fading
in after the character speaks, building gently toward the end.

Or silence:

No background music. Ambient sound only.

Multi-Shot Prompting with H3

H3 supports multi-shot generation within a single prompt. Use [Shot N] labels and timestamp-based cuts.

Format

[Shot 1] Opening shot description. Duration and action.

[Shot 2] At 00:03.000, cut to next shot. Description.

[Shot 3] At 00:07.000, cut to final shot. Description.

Example: 10-Second Product Ad (3 Shots)

[Shot 1] Real-life ad, extreme close-up of a runner tying
shoelaces on a wet track at pre-dawn. Camera low and static.
Fabric tightens, lace pulls taut. 3 seconds.

[Shot 2] At 00:03.000, cut to side tracking shot as the runner
accelerates through the first curve. Each footstep splashes
a small spray from the track. Breathing and foot impact
overtake distant city ambience. 4 seconds.

[Shot 3] At 00:07.000, cut to the shoe landing beside the
finish line, then slow push-in as the runner stops. Brief
bass pulse ends on the final footstep. 3 seconds.

Multi-Shot Best Practices

  1. Allocate time deliberately — 3s for setup, 4s for action, 3s for ending
  2. Keep each shot simple — one action, one camera move per shot
  3. Use timestamps for cuts, not for continuous shots
  4. Maintain identity — reuse exact character descriptions across shots
  5. Audio continuity — describe how sound transitions between shots
  6. Don’t timestamp a single continuous shot — use timestamps only for cuts between shots

25 Copy-Ready MiniMax H3 Prompt Templates

Product & Commercial (5 Templates)

1. Product Hero with Liquid

Real-life studio product video. A [product] stands on a
[ surface]. A narrow stream of [liquid] strikes the surface
beside it, splashes curling around the base. Camera orbits
slowly clockwise with the label facing the lens. Hard edge
light, restrained reflections. The shot ends with the liquid
settled and the product centered.

2. Product Reveal

[Shot 1] Close-up of a branded box on a dark surface.
Hands slowly open the lid, revealing [product] inside.
Soft overhead light fills the box. 4 seconds.

[Shot 2] Camera slowly pulls back to reveal the full product.
Hands lift it out and turn it to face camera. Clean studio
lighting. 6 seconds.

3. Text Overlay Commercial

Real-life product commercial. [Product] on [surface].
Animated text "[BRAND TEXT]" slides in from below and locks
into position with a subtle bounce. Camera static. Clean
modern design. Motion graphics style. No text distortion.

4. Watch Macro Detail

Extreme macro of a luxury watch face on dark marble. The
second hand sweeps smoothly. Hard pin-spot from above
catches the polished bezel. Camera slowly pushes in toward
the dial center. Tick-tock sound, no music. 8 seconds.

5. Jewelry Rotation

A diamond ring on a black velvet pedestal rotates slowly
on a turntable. Single hard spotlight from upper right
catches each facet. Camera static close-up. Deep black
and cool white sparkle palette. 7 seconds.

Character & Dialogue (5 Templates)

6. Character Walking (Physics-Accurate)

Low-angle tracking shot at street level. A [character
description] walks through [location]. Steady pace. Each
step lands heel-first, rolling forward with visible weight
transfer. [Surface] reflects [lighting]. 35mm film, shallow
depth of field. Ambient [sound].

7. Two-Person Dialogue

[Shot 1] Medium close-up of [Character A] facing camera.
[Character A: description, voice]: "[dialogue]." Natural
[lighting]. 4 seconds.

[Shot 2] Cut to [Character B] responding.
[Character B: description, voice]: "[dialogue]." Same
lighting setup. 4 seconds.

8. Emotional Close-Up

Extreme close-up of [character]'s face. [Emotion] crosses
their features slowly—eyes soften, shoulders drop slightly.
Single hard key from camera-[direction]. Shallow depth of
field. Intimate, vulnerable. No face morphing. 6 seconds.

9. Character Interaction with Object

[Character] reaches out and picks up [object] from [surface].
Fingers wrap around it with natural grip. Camera slowly
pushes in. [Lighting]. Cinematic [style]. No hand deformation,
no object distortion. 6 seconds.

10. Character with Wind

[Character] stands at [location]. Gentle wind from camera-
[right] lifts and flows their [hair/clothing] naturally.
Camera static medium shot. Golden hour backlight. Cinematic.
No hair merging, no face distortion. 6 seconds.

UGC & Social Media (5 Templates)

11. Creator Talking Head

Vertical phone video. Creator faces camera in a bright room.
"[Hook line]" in an enthusiastic, genuine tone. Natural
window light. Slight handheld movement. Social-first pacing.
No face distortion. 8 seconds.

12. Unboxing

Overhead shot. Hands open a branded box on a table, lift
out [product], and turn it to show the camera. Natural
indoor lighting. UGC style. "Look what just arrived" in an
excited whisper. No hand clipping. 8 seconds.

13. Before/After Reveal

[Shot 1] [Before state]. Camera static. Clean background.
3 seconds.

[Shot 2] Quick cut to [after state]. Same framing. Bright
studio lighting. Satisfying transformation moment. 4 seconds.

14. Product Demo

Creator holds [product] to camera. Demonstrates [feature]
with natural hand motion. "[Product demo line]" in a
confident, helpful tone. Warm studio lighting. Social
pacing. No hand deformation. 10 seconds.

15. Flat Lay Animation

Overhead shot of a styled flat lay arrangement. One element
slowly slides into frame from the edge and settles into
position. Camera static. Soft even daylight. Editorial
style. No element overlap. 6 seconds.

Cinematic & Nature (5 Templates)

16. Establishing Shot

Wide establishing shot of [location] at [time of day].
[Atmospheric element] drifts across the scene. Camera
slowly cranes up revealing the full scale. Cinematic [color]
palette, volumetric light. Ambient [sound], no dialogue.
10 seconds.

17. Storm Sequence

[Shot 1] Wide shot of [location] as storm approaches.
Darkening sky, wind picks up. 4 seconds.

[Shot 2] Rain begins, lightning flash illuminates the scene.
Camera static. Deep thunder. 4 seconds.

[Shot 3] Downpour. Camera slowly pushes in toward [subject].
Rain on surfaces. 4 seconds.

18. Golden Hour Landscape

Wide shot of [landscape] at golden hour. Long shadows stretch
across the terrain. Warm low sun. Camera slow crane up
revealing the full vista. Gentle wind moves vegetation.
Ambient natural sound. No music. 10 seconds.

19. Foggy Forest Walk

Medium wide shot moving slowly through a dense forest at
dawn. Thick fog drifts between trees. Volumetric light rays
through the canopy. Camera tracks forward at walking pace.
Muted green and grey palette. Soft ambient forest sounds.
10 seconds.

20. Underwater Scene

Wide shot underwater. [Subject] moves through the water
with natural buoyancy. Dappled light from the surface
creates moving caustics. Camera slowly tracks alongside.
Dreamy pace. Aquatic ambient sound. 8 seconds.

Food & Beverage (5 Templates)

21. Coffee Pour

Close-up of a ceramic cup on a wooden surface. Coffee pours
from above in a smooth stream, filling the cup. Steam rises
slowly. Camera slowly pushes in. Warm café lighting, soft
bokeh background. Café ambience. 6 seconds.

22. Sizzling Dish

Overhead close-up of [dish] on a hot plate. [Ingredient]
sizzles and pops. Steam rises. Camera slowly pushes in.
Warm kitchen lighting. Sizzling sound, no music. 6 seconds.

23. Cocktail Assembly

Close-up of a glass on a bar surface. [Liquid] pours over
ice in slow motion. Ice cracks and shifts. Camera slowly
orbits. Dramatic bar lighting. Liquid sound, clinking ice.
No liquid clipping. 7 seconds.

24. Ingredient Prep

Close-up of hands slicing [ingredient] on a wooden cutting
board. Each cut is precise and rhythmic. Overhead angle.
Bright natural kitchen lighting. Knife-on-wood sound.
No hand deformation. 6 seconds.

25. Chocolate Break

Close-up of a chocolate bar. A hand slowly breaks off a
piece, revealing the cross-section. The piece separates
cleanly with a satisfying snap. Camera static macro. Warm
studio lighting. Snap sound, no music. 5 seconds.

For more tested H3 prompts with their generated outputs, browse the VideosPrompt community which features prompts organized by model and category.


Troubleshooting: When H3 Ignores Your Prompt

When output fails, change the minimum instruction that explains the failure. Don’t rewrite the entire prompt.

Failure Likely Cause First Fix
Random cuts Multiple camera moves or scene changes competing Request one continuous shot, one camera path
Wrong person speaking Speaker identification inconsistent Assign stable labels to each speaker, shorten dialogue
Dialogue feels rushed Lines too long for the duration Time lines aloud, or increase clip duration
Product deforms Too many actions, camera, and style elements Lock product orientation, remove secondary motion
Reference identity drift Conflicting references or unclear character Keep the clearest identity image, state remaining characters explicitly
First and last frame don’t connect Missing physical transition Describe the intermediate poses, objects, and camera changes
Audio feels unrelated Audio described as mood only Name the source, timing, and volume changes
Subject ignores camera Camera instruction buried in description Move camera instruction to its own sentence, early in the prompt
Style inconsistency Multiple style descriptors fighting Use one style label: “cinematic” or “documentary,” not both

Key principle: If the camera is wrong but the subject is correct, keep the subject language and modify only the camera sentence. Each retry should change one variable.


H3 vs. Competitors: Prompt Style Differences

The same concept needs different prompt language for each model:

Same Scene, Four Dialects

Concept: A woman walks through a rain-soaked alley at night.

MiniMax H3:

Real-life cinematic shot. A woman in a dark coat walks
through a rain-soaked alley at night. Steady pace, each
step splashes in puddles. Camera tracks alongside at waist
height. Neon reflections on wet pavement. Ambient rain,
distant traffic. She stops at the end of the alley and
looks up. No music.

Kling 3.0:

Medium tracking shot. Woman in dark coat walks through
rain-soaked alley at night. Steady pace, puddle splashes.
Neon reflections on wet pavement. 35mm film, shallow DOF.
Ambient rain, distant traffic.

Seedance 2.0:

Woman in dark coat walking through rain-soaked neon alley
at night. Side tracking shot, steady pace. Wet pavement
reflections. Cinematic, moody. Rain ambience, no music.

Sora 2:

A woman in a dark overcoat walks steadily through a narrow
rain-soaked alley at night. Neon signs reflect off the wet
pavement in pools of color. The camera follows her from the
side at waist height. Rain falls gently. Cinematic, moody
atmosphere.

H3 prompts read like timed shot lists. Kling reads like structured controls. Seedance reads like scene briefs. Sora reads like descriptive prose.


MiniMax H3 Version Comparison

Version Resolution Duration Audio Key Feature Best For
H3 (Base) Up to 2K 5-15s Stereo Open weights, multi-modal Local deployment, experimentation
H3 Max (fal) Up to 2K 5-15s Enhanced stereo Post-trained, Director endpoint Production quality, commercial
H3 Turbo 768p 5-10s Stereo Fast generation Drafts, batch testing

When to use each:

  • H3 Base: Open-source workflows, ComfyUI, local GPU deployment
  • H3 Max: Production commercials, brand films, hero content
  • H3 Turbo: Rapid iteration, A/B testing, draft previews

FAQ: MiniMax H3 Prompting

What is the MiniMax H3 prompt formula?

The 6-block formula: [Opening Frame] + [Subject Action] + [Environment Response] + [Camera Path] + [Audio] + [Final Frame]. This sequence gives H3 a timeline to follow rather than a pile of adjectives to interpret.

How long should a MiniMax H3 prompt be?

H3 accepts up to 7,000 characters, but most effective prompts are 3-6 sentences. Longer isn’t better—clearer is better. Each sentence should specify one dimension: opening frame, action, camera, audio, or final state.

How do I write dialogue for MiniMax H3?

Put exact lines in quotes. Label each speaker with description and voice tone. Keep lines short—time them aloud. Specify language. For multi-speaker scenes, add “their lips only move during their own lines” to prevent cross-sync issues.

What’s the difference between H3 text-to-video and image-to-video?

Text-to-video requires you to define the entire opening frame from scratch. Image-to-video uses your uploaded image as the opening frame—your prompt should only describe what happens AFTER the image, not re-describe what’s already visible.

How many reference files can H3 accept?

Up to 12 files total: 9 images, 3 video clips (2-15 seconds each), and 3 audio clips. Each reference should have an explicit role—don’t upload conflicting references that ask the model to guess which to follow.

Can H3 render text and UI elements?

Yes. H3 is notably strong at rendering legible text, subtitles, brand logos, and UI elements. Specify the exact text you want rendered and its position on screen.

Why does H3 ignore parts of my prompt?

Common causes: too many stacked actions, conflicting camera moves, ambiguous reference characters, or dialogue that’s too long for the duration. Remove secondary instructions and test one change at a time.

Where can I use MiniMax H3?

H3 is available on fal.ai (API, pay-per-use), C Dance AI (web interface), ComfyUI (local deployment with open weights), and the MiniMax platform. The VideosPrompt community also features H3 prompts you can study and adapt.


Conclusion

MiniMax H3 is a director’s model. It doesn’t want adjectives—it wants a shot list.

The framework is simple:

  1. Define the opening frame — what exists before motion starts
  2. Describe one action — with cause and effect, not just the verb
  3. Set the camera — one move, one path, one instruction
  4. Direct the sound — dialogue, ambience, effects, music as separate layers
  5. Anchor the final frame — where things end up
  6. Protect fragile elements — explicit constraints for text, hands, faces

Start with the templates above. Swap in your subjects. Generate, compare, refine.

Every clip builds your intuition for how H3 interprets timed instructions. After 50 generations, you won’t need templates—you’ll look at a scene and see the shot list underneath.

Your stories deserve 2K resolution and stereo sound. Now you know how to direct them.


Last updated: September 2026

Ready to explore more? Browse the VideosPrompt community for thousands of curated AI video prompts across every model—including dedicated MiniMax H3 collections. Or try SocialToPrompt to reverse-engineer any video into a reusable prompt.

Share Article

Related Articles

Recommended Reading

Ready to Get Started?

Experience our product immediately and explore more possibilities.