Multimodal Reference Video Prompts: Combining Images, Videos, and Audio as Input
TL;DR
- Multimodal reference video prompts route text, image, video, and audio inputs together so each input owns a distinct aspect of the final render (subject, motion, ambience, beat).
- Single-mode prompts plateau once you need character lock + camera path + ambient sound + visual style simultaneously; multi-input composition is the only way to keep all four under control.
- The three primary input modes are image (subject, style, palette), video (motion, choreography, camera move), and audio (rhythm, ambience, voice, foley). Each must be assigned a role in the prompt.
- The hardest part is not the upload — it is routing: telling the model which input owns the subject, which owns motion, which owns lighting, so references do not fight each other.
- Model support matrix in 2026: Seedance 2.5 supports up to 50 asset references, Veo 3.1 ingests image + audio, Kling 3.0 uses elements slots, and Wan 3.0 supports start-end frames plus audio.
- Cross-modal anchoring (explicit role assignment per reference) is what makes character consistency + camera choreography work together; without it, the encoder picks one reference and ignores the rest.
- Use negative-prompt baselines specific to multimodal (“no style bleed between image and video references”, “no subject shift across shots”) to suppress the common failure modes.
What “multimodal” actually means in 2026 video generation
A multimodal reference video prompt is one in which the model receives more than one type of input signal — typically a text prompt, one or more images, one or more reference videos, and optionally an audio file — and the prompt explicitly states what each input is responsible for.
This sounds simple. In practice, it is the largest usability shift in video generation since image conditioning landed in 2024. Up until late 2025, most production video workflows were either text-only or text-plus-a-single-image. That ceiling showed up everywhere: you could lock a face from one image, but you could not simultaneously lock the camera move from a reference clip; you could describe ambient cafe sound, but you could not pin it to an actual audio file; you could not keep a character’s wardrobe consistent across a 12-shot sequence because every regeneration re-rolled the costume.
The 2026 generation stack — Seedance 2.5, Veo 3.1, Kling 3.0, Wan 3.0 — breaks that ceiling by accepting multiple references in parallel and letting each one carry an independent signal. But the model does not decide for you which input means what. If you feed it a portrait of a woman and a still from a film, and write “match this aesthetic”, the model will pick whichever reference reads loudest in its encoder and ignore the other. The prompt’s job is to disambiguate that.
Three trends are converging to make multimodal the default rather than the exception. First, reference-image support is now table stakes across every major model — even the lower-tier players ingest at least one image. Second, reference-video slots have moved from experimental to production-grade: Seedance 2.5’s ref2v path and Veo 3.1’s reference-video slots both ship as stable APIs, not research previews. Third, audio conditioning has crossed the threshold from a novelty into something you can build real scenes around, particularly for Veo 3.1’s ambient-and-voice pipeline and Wan 3.0’s rhythm-anchor slot.
The cost of all this is prompt complexity. A text-only prompt can be a single sentence. A multimodal prompt is closer to a shot list with attached metadata: which input is the subject, which is the lighting, which is the camera, which is the beat. That is what the rest of this article teaches.
This is the entire discipline of multimodal reference prompting. The rest of this article covers the input modes, the routing problem, the model-by-model support matrix, anchoring techniques, and 12 working templates.
For background on single-input patterns, see our Seedance 2.5 prompt guide and the companion piece on Seedance multi-reference image prompts, which covers the same routing discipline for image-only references.
The three input modes
Every multimodal video prompt is a combination of three input modes (plus text). Understanding each mode’s native signal is prerequisite to routing them correctly.
1. Image input — subject, composition, and style
An image reference contributes three things to a generation: the subject (face, body, wardrobe, props), the composition (framing, depth, foreground/background layering), and the style (color grading, lighting direction, texture, era). Models that ingest image references do so by encoding them into a vision-language embedding; the text prompt then steers which of those three facets the embedding should dominate.
A practical pattern:
Image #1 (portrait.jpg): woman, late 20s, red trench coat, soft window light.
Text: "Use Image #1 as the subject reference only. Do not copy its lighting. Place her in a rainy neon alley, mid-stride, three-quarter framing."
Here the image is routed to “subject only”, and the text explicitly overrides its lighting. Without that override, the model will copy the image’s full mood and your text prompt becomes decorative.
Image references are also the cleanest way to lock character consistency across multiple shots — for a deeper treatment of that workflow, see AI character consistency prompts 2026 and image-to-video character lock prompts.
2. Video input — motion, camera path, choreography
A video reference contributes temporal information that no image can: camera move, character motion, scene pacing, lens behavior, and the cadence of cuts. When you upload a reference clip, the model encodes its motion vectors and uses them as a prior.
Video #1 (whip_pan.mp4): handheld dolly shot panning across a crowded market.
Text: "Match the camera movement and rhythm of Video #1. Subject is a chef in white apron (described in text). Do not copy the market environment — render a quiet kitchen instead."
The pitfall here is over-copying. A video reference carries environment, subject, lighting, AND motion in its pixels. If your text prompt says nothing about which facets to keep, the model typically defaults to motion and discards environment — but in some models the environment wins. Always disambiguate.
For models that specialize in video-as-input (Seedance 2.5’s ref2v and Veo 3.1’s reference-video slots), this is the most powerful single mode in the toolkit. The Ref2V companion article walks through the mechanics in detail.
3. Audio input — rhythm, ambience, voice, beat sync
Audio input is the newest of the three and the least uniformly supported. As of October 2026, only Veo 3.1 and Wan 3.0 ingest audio files natively (Seedance 2.5 added an experimental audio-conditioning slot in its 2.5 update, and Kling 3.0 uses it for timing only, not content). Audio contributes four signals: ambient sound bed (cafe, wind, room tone), rhythm/beat (which the model can sync motion to), dialogue/voice (which some models lip-sync), and musical cues (which can drive transitions).
Audio #1 (rain_loop.wav): steady rainfall on a tin roof, no music.
Text: "Generate 8 seconds of a man walking through a forest at night. Match the rhythm of Audio #1's rainfall cadence — slow footsteps in sync with each drip cluster. No dialogue."
This is the cleanest example of multimodal working: a single audio file is doing rhythmic anchoring that would be impossible to specify in words.
For Veo 3.1’s full audio stack, see Veo 3.1 native audio 4K.
Model support matrix (October 2026)
Not every model ingests every mode. The matrix below summarizes what each of the four major video models in our testing accepts as input, and what each input controls.
| Model | Image refs | Video refs | Audio input | Max references | Routing primitive |
|---|---|---|---|---|---|
| Seedance 2.5 | Yes | Yes | Experimental (beat only) | Up to 50 asset references across all modes | Per-asset role tag ([subject], [motion], [style]) |
| Veo 3.1 | Yes | Yes | Yes (ambient + voice + beat) | 3 images + 2 videos + 1 audio (typical) | Reference slot ordering |
| Kling 3.0 | Yes (elements slots) | No native video ref | No | Up to 4 element slots | Element slot assignment |
| Wan 3.0 | Yes (start + end frame) | Implicit (via keyframe interpolation) | Yes (ambient + beat) | 2 frames + 1 audio | Start/end keyframe + audio conditioning |
Sources for the matrix above: Seedance 2.5 documentation as summarized by mindstudio.ai’s Seedance 2.5 features breakdown and seedance2ai.net’s 50-references guide; Veo 3.1 capabilities per DeepMind’s official Veo model page and the Veo API cookbook. Kling 3.0 and Wan 3.0 rows reflect observed behavior in the public-facing tooling as of October 2026.
Two things stand out from this matrix. First, Seedance 2.5 is the only model in the group that scales reference count past a handful — the 50-asset cap is unusual and unlocks complex multi-element scenes. Second, Veo 3.1 is the only model with a clean, production-grade audio slot that ingests real sound files rather than just rhythm metadata.
The routing problem
The single hardest thing about multimodal reference video prompts is not the upload step. It is routing: deciding which input owns which aspect of the output, and stating that decision in the prompt.
The naive failure mode looks like this: you upload a portrait of your character AND a still from a Blade Runner film. You write “cyberpunk woman walking through neon market”. The model produces something, but it is unclear whether the woman’s face came from your portrait or whether the lighting came from your reference still. You regenerate four times and get four different answers.
What is happening: the model is treating both references as equally-weighted style priors. Without an explicit role assignment, it samples one or the other at random per generation.
The fix is to give each reference a role tag in the prompt. The exact syntax varies by model:
- Seedance 2.5: use bracket tags like
[subject=portrait.jpg],[lighting=blade_runner_still.jpg],[motion=walking_loop.mp4]. The model weights each slot independently. - Veo 3.1: references are uploaded into ordered slots. The first image slot is treated as primary subject; subsequent slots are weighted in order. The text prompt can explicitly demote a slot (“Image #2 is style-only — ignore its composition”).
- Kling 3.0: each element slot is a labeled entity. You tag the slot (“character”, “outfit”, “prop”) when uploading.
- Wan 3.0: start frame and end frame are unambiguous roles by default; audio slot is also unambiguous.
If you skip the role assignment on a model that does not enforce it (Kling, Wan), you get random sampling. On a model that does enforce it (Seedance, Veo), the unassigned slot defaults to “style” — which may or may not be what you wanted.
A reliable heuristic: write the role assignment in the prompt BEFORE describing the scene. A prompt that opens with “Use Image #1 as subject reference, Image #2 as lighting reference, Video #1 as camera motion” will route correctly on every model that supports the inputs, regardless of slot ordering.
There is also a secondary failure mode worth naming. Even when role assignment is correct, models occasionally drift: an image tagged as subject-only will, on a bad regeneration, leak its lighting into the final frame, or a video tagged as motion-only will bleed its color grading into the output. This is not a routing failure in the prompt — it is an encoder failure inside the model. The mitigation is twofold. First, write tight negative prompts (see the section below). Second, run several regenerations per shot and cherry-pick the cleanest one. The multimodal workflow is not deterministic in the way a text-only prompt can occasionally be; budget for three to five regenerations per usable clip on average.
Cross-modal anchoring techniques
Cross-modal anchoring is the set of techniques that let a single reference play multiple roles across the generation, or that lock two different references to two different facets of the same shot.
Technique 1: Video-as-motion-template, image-as-subject
The most common anchoring pattern. You have a reference video showing the camera move you want, and a portrait of the character who should appear in the shot. Without anchoring, the model will substitute the character from the video. With anchoring, you tell the model:
Video #1 [motion only]: handheld dolly-in from wide to medium close-up over 4 seconds.
Image #1 [subject only]: woman in red trench coat.
Text: "Render Video #1's camera move. Place Image #1's character in the frame. Background: empty warehouse."
This is exactly the pattern Seedance 2.5 was designed for — the Ref2V workflow is built on it.
Technique 2: Image-as-style, video-as-pacing
The reverse: you want the look of an image (a Caravaggio painting) but the pacing of a reference video (a long, slow Derek Cianfrance shot). Without explicit role tagging, the model will copy the image literally (a static painting) and discard the video entirely.
Image #1 [style only]: Caravaggio's "The Calling of Saint Matthew", use only its color palette and chiaroscuro lighting.
Video #1 [motion only]: 8-second slow push-in with no cuts.
Text: "A modern executive sits at a desk. Light her as Image #1. Move the camera as Video #1."
Technique 3: Audio-as-rhythm-anchor, image-as-style-and-subject
For Veo 3.1 and Wan 3.0, you can pin a scene’s pacing to an audio file’s beat while keeping visual style and character from image references:
Image #1 [subject + style]: young man in 1970s suit, saturated Kodachrome palette.
Audio #1 [rhythm only]: drum loop at 92 BPM.
Text: "Image #1's character walks across a motel parking lot at night. Match his footstep cadence to Audio #1's kick drum. No dialogue."
Technique 4: Stacked-element anchoring (Seedance 2.5 only)
Seedance 2.5’s 50-reference slot lets you assign different roles to many small references at once. The Seedance multi-reference image prompts article covers this in depth — the short version is you can pin separate references for the character’s face, the character’s outfit, the prop, the background, the lighting direction, and the camera move all in one generation.
12 prompt templates
The templates below are written to be drop-in: copy the prompt, supply the referenced files, and run. Each is labeled with the model it was tested on primarily; most work across models with minor syntax changes.
Text + 1 image (3 templates)
Template 1 — Character lock from portrait
Model: Veo 3.1 (also runs on Seedance 2.5, Kling 3.0)
Input: 1 image (character_portrait.jpg)
Role: Image = subject only. Text owns everything else.
[Reference: character_portrait.jpg] Subject reference only.
A woman in her late 20s walks through a rain-soaked Tokyo alley at night.
She wears a red trench coat over a white shirt, as in the reference.
Camera: handheld, slightly low angle, follows her from the side.
Lighting: neon signage from above, puddles reflecting pink and teal.
Duration: 6 seconds. Aspect 16:9. No dialogue. No cuts.
Template 2 — Style transfer from painting
Model: Wan 3.0 (also runs on Veo 3.1)
Input: 1 image (oil_painting.jpg)
Role: Image = color palette and texture only. Text owns composition and motion.
[Reference: oil_painting.jpg] Use only its color grading and brushstroke texture.
A chef in a white apron plates a dish in a quiet kitchen.
Adopt the reference's muted ochre-and-slate palette and visible canvas grain.
Camera: static medium shot, gentle rack focus from hands to face.
Duration: 5 seconds. Ambient sound only.
Template 3 — Product shot from hero image
Model: Seedance 2.5
Input: 1 image (product_hero.jpg)
Role: Image = subject + composition. Text adds motion and lighting.
[Reference: product_hero.jpg] Subject and framing reference.
Render the bottle from the reference on a rotating turntable.
Slow 360-degree orbit over 8 seconds.
Background: deep navy gradient, single soft top light catching the glass.
No text overlays. No hands.
Text + 1 video (3 templates)
Template 4 — Camera path reuse
Model: Seedance 2.5
Input: 1 video (dolly_clip.mp4)
Role: Video = camera motion only. Text replaces subject and environment.
[Reference: dolly_clip.mp4] Camera motion template only.
Match the exact dolly-in trajectory and speed of the reference.
Subject: a young boy running through a wheat field at golden hour.
Replace the reference's environment entirely.
Duration: 5 seconds. No dialogue.
Template 5 — Pacing and rhythm reuse
Model: Veo 3.1
Input: 1 video (long_take.mp4)
Role: Video = pacing only (no environment, no subject).
[Reference: long_take.mp4] Use only the pacing — single continuous shot, no cuts, slow acceleration over 6 seconds.
Subject: a chess match between two elderly men in a park.
Do not copy the reference's setting.
Lens: 50mm equivalent, eye level.
Sound: park ambience, distant traffic.
Template 6 — Choreography reuse
Model: Seedance 2.5
Input: 1 video (dance_clip.mp4)
Role: Video = choreography only. Subject and wardrobe from text.
[Reference: dance_clip.mp4] Choreography reference only.
Copy the dance sequence beat-for-beat.
Two dancers in matching black outfits perform the routine on an empty soundstage.
Camera: wide static shot to capture full body.
Lighting: single overhead spotlight, hard shadows.
Duration: 8 seconds. No music in audio (we will score in post).
Text + image + video (3 templates)
Template 7 — Subject from image, motion from video
Model: Seedance 2.5 (canonical use case)
Input: 1 image (character.jpg), 1 video (walk_loop.mp4)
Role: Image = subject. Video = motion. Text owns environment and lighting.
[Reference: character.jpg] Subject only — face, hair, wardrobe.
[Reference: walk_loop.mp4] Motion only — match the walk cycle and arm swing exactly.
Place the subject in a crowded train station at rush hour.
Camera: tracking shot from the front, slight low angle.
Lighting: fluorescent overhead, mixed with daylight from entrance.
Duration: 6 seconds. Ambient sound: station PA, footsteps, distant train.
Template 8 — Style from image, pacing from video, subject from text
Model: Veo 3.1
Input: 1 image (film_still.jpg), 1 video (long_lens.mp4)
Role: Image = color grading and lens choice. Video = pacing. Text owns subject.
[Reference: film_still.jpg] Color grading and lens only — adopt its desaturated teal-and-orange palette and shallow depth of field.
[Reference: long_lens.mp4] Pacing only — single 8-second shot, no cuts, slow lateral push.
Subject: a journalist interviewing a source in a dimly lit bar.
Do not copy the reference video's setting or subjects.
Template 9 — Multiple image refs (style + subject) + video (camera)
Model: Seedance 2.5
Input: 2 images (face.jpg, wardrobe.jpg), 1 video (camera_move.mp4)
Role: face = subject identity. wardrobe = costume. video = camera move.
[Reference: face.jpg] Subject's face only.
[Reference: wardrobe.jpg] Wardrobe and accessories only — red dress, gold earrings, as shown.
[Reference: camera_move.mp4] Camera trajectory only — match the slow crane-up.
The subject walks through an art gallery, pausing at a painting.
Lighting: warm gallery spots, neutral walls.
Duration: 7 seconds.
Text + image + video + audio (3 templates)
Template 10 — Full multimodal: character, motion, audio ambience
Model: Veo 3.1 (canonical four-mode prompt)
Input: 1 image (char.jpg), 1 video (pace.mp4), 1 audio (rain.wav)
Role: image = subject, video = pacing, audio = ambient bed and rhythm anchor.
[Reference: char.jpg] Subject only — woman in her 50s, grey hair, beige coat.
[Reference: pace.mp4] Pacing only — match the slow lateral track over 6 seconds.
[Reference: rain.wav] Ambient sound and rhythm anchor — match her walking cadence to the rainfall's intensity curve.
She walks through a park at dusk during steady rain.
No dialogue. No music.
Template 11 — Music video shot with locked choreography
Model: Seedance 2.5 (with experimental audio)
Input: 1 image (performer.jpg), 1 video (choreo.mp4), 1 audio (track.wav)
Role: image = performer identity, video = choreography, audio = beat-sync and lyrics.
[Reference: performer.jpg] Subject only — keep the performer's face and signature outfit.
[Reference: choreo.mp4] Choreography only — copy the routine's structure and timing.
[Reference: track.wav] Beat-sync and lip-sync — match movement to kick drum on beats 1 and 3.
Lip-sync to the chorus's vocal melody.
Background: dark soundstage with three rotating spotlights.
Duration: 12 seconds. Single shot.
Template 12 — Documentary interview with locked look and ambient sound
Model: Veo 3.1
Input: 1 image (subject.jpg), 1 video (interview_style.mp4), 1 audio (room_tone.wav)
Role: image = subject, video = interview framing and gesture pacing, audio = room tone.
[Reference: subject.jpg] Subject identity only — older man with glasses, navy sweater.
[Reference: interview_style.mp4] Framing and gesture pacing only — medium close-up, subject nods on the third beat of each phrase.
[Reference: room_tone.wav] Room tone only — quiet indoor ambient with occasional HVAC hum.
The subject sits in a leather chair, speaking directly to camera about his career.
No background music.
Duration: 8 seconds.
Negative-prompt baselines for multimodal
Single-mode video generation has its own negative-prompt vocabulary. Multimodal generation has an additional layer, because references can conflict with each other in specific, recurring ways. The negative baselines below are the ones we use across every multimodal prompt:
Universal multimodal negatives:
- no style bleed between image and video references
- no subject shift across shots (character must remain identical to Image #1 in every frame)
- no lighting conflict between reference and text description
- no environment copy from video reference (unless explicitly assigned)
- no costume drift between reference image and generated frames
- no camera move that contradicts the video reference's trajectory
- no audio sync drift between motion and rhythm input
Add these to your prompt’s negative section for every multi-input generation. They address the most common failure modes — references fighting each other for control of an aspect.
Per-model additions:
- Seedance 2.5: no reference priority inversion (when one slot bleeds into another’s role).
- Veo 3.1: no off-slot reference contamination (when image #2 affects subject instead of style).
- Kling 3.0: no element slot bleed (when prop slot affects character).
- Wan 3.0: no start-end frame averaging (when model blurs the keyframes instead of interpolating).
FAQ
What is a multimodal reference video prompt? A prompt that combines multiple input types — typically text plus one or more images, reference videos, and/or audio files — with explicit instructions about what each input controls in the generated video.
Which 2026 models support multimodal reference inputs? Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0 all accept some combination of image, video, and audio references. Seedance 2.5 has the largest reference budget (up to 50 assets); Veo 3.1 has the most complete audio integration; Kling 3.0 uses element slots for image-only composition; Wan 3.0 uses start/end keyframes plus audio conditioning.
How do I prevent references from conflicting? Assign each reference an explicit role in the prompt (subject, lighting, motion, style, pacing). If the model supports role tags, use them. If not, write the role assignment in plain text before describing the scene. See the routing problem section above.
Can I use a video reference for both motion and environment? Technically yes, but you should not. If you want both, supply two references — one tagged motion-only, one tagged environment-only. Models handle single-role references far more reliably than multi-role ones.
Do audio references work for voice and dialogue? Only in Veo 3.1’s voice slot, and only for short utterances. Lip-sync is not reliable across most models as of October 2026. For dialogue-heavy scenes, generate silent and dub in post.
How many references can I use at once? Depends on the model. Seedance 2.5 caps at 50 asset references; Veo 3.1 typically accepts 3 images + 2 videos + 1 audio; Kling 3.0 caps at 4 element slots; Wan 3.0 is start/end + audio. More references do not always mean better output — most production prompts use 2-4 references and rely on text for the rest.
What is the most common multimodal prompting mistake? Forgetting to assign roles. Uploading references without telling the model which aspect each one controls leads to random sampling and inconsistent regenerations. Always write the role assignment first.
Conclusion
Multimodal reference video prompting is the discipline of routing — telling the model which input owns which facet of the output. Once you internalize that, the rest is mechanical: choose your references, assign each one a role, describe the scene, and add the right negative baselines.
The 12 templates above are starting points. The model landscape will keep shifting — Seedance 2.5’s 50-asset slot will likely be matched by competitors in 2027, audio conditioning will become universal, and video-to-video references will replace some image references. The routing discipline stays constant.
For deeper reading on adjacent workflows:
- Seedance multi-reference image prompts — pure image-reference composition at scale
- Seedance 2.5 prompts — model guide and base prompt patterns
- Veo 3.1 native audio 4K — Veo’s audio stack in detail
- Ref2V companion article — video-as-input deep dive
- AI character consistency prompts 2026 — keeping characters stable across shots
- Image-to-video character lock prompts — first-frame character anchoring
External references used in this article:
- Seedance 2.5 features breakdown — mindstudio.ai
- Seedance 2.5 50-references guide — seedance2ai.net
- Veo model page — DeepMind
- Veo cookbook — Google AI Studio
Reviewed by the videosprompt.org editorial team · October 2026
Share Article