Reference-to-Video Templates: A Cross-Model Playbook for Seedance, Veo, Kling, Wan, and Beyond
TL;DR
- Reference-to-video (Ref2V) uses one or more still images as the visual anchor for a generated video clip, rather than relying on text alone.
- The three core Ref2V techniques are single-subject reference, multi-subject reference, and style-only reference, each with distinct scaffolds and failure modes.
- This article ships 12 copy-ready templates adapted for Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0.
- A universal Ref2V skeleton — anchor, subject token, scene, motion/camera, audio — travels across all four model APIs.
- MiniMax H3 is referenced but unverified. We have not run H3 against these templates; the playbooks are validated against the four documented APIs with a checklist for H3.
- Style-only references are the cheapest to validate and the best starting point for any new model including H3.
- A verification protocol at the end lets you score any model, including H3, against the same five-component skeleton.
1. What “reference-to-video” means in 2026
Reference-to-video (Ref2V) conditions a video model on one or more input images so the output inherits specific visual properties — a face, a product, a lighting style, a brand palette — that text alone cannot reliably describe. The Ref2V paper (arXiv:2508.02458) formalized the subject-token handoff pattern that all major commercial APIs now expose under different names.
Three core techniques sit inside Ref2V:
| Technique | Primary use case | Inputs | Failure mode |
|---|---|---|---|
| Single-subject reference | Lock one identity (face, product, character) across a clip | 1 reference image + scene description | Identity drift when scene prompt dominates the reference |
| Multi-subject reference | Composite two or more characters/objects into one scene | 2–N reference images + role tags | Subject bleed and pose confusion when tokens are not disambiguated |
| Style-only reference | Inherit lighting, palette, or artistic treatment | 1 reference image + style-only prompt | Style overpowers subject; subjects get re-stylized rather than re-rendered |
The distinction matters because each technique has a different prompt skeleton, cost profile, and risk of failure. A practical Ref2V playbook handles all three without collapsing them.
2. What we verified, what we couldn’t
Verified against documented APIs. The 12 templates are adapted against publicly documented surfaces of four models:
- Seedance 2.5 —
@tagreference syntax and 50-asset workflow (Seedance 50 references guide). - Veo 3.1 —
image_promptfield (deepmind.google/models/veo). - Kling 3.0 —
elementsarray for multi-reference binding. - Wan 3.0 —
start_imageandend_imagefields.
For prompt engineering context, see the mstudio.ai filmmaking prompting guide.
MiniMax H3 is referenced but unverified. The keyword for this piece names H3, so we include it in the title for discoverability. We have not run H3 against these templates and do not know whether H3 uses @tag-style syntax, image_prompt, elements[], or another mechanism. Apply these to H3 as a hypothesis and use Section 8 to score actual output.
We are not going to fabricate H3-specific field names or claim benchmark numbers we did not run.
3. The Ref2V prompt skeleton
Every Ref2V prompt decomposes into five components. We use this skeleton across all 12 templates because it survives the translation between APIs better than model-native syntax does.
| # | Component | Purpose | Universal example |
|---|---|---|---|
| 1 | Source reference anchor | Tells the model which image(s) to consult | reference_image: ./subject.png |
| 2 | Subject token description | Names what the anchor represents so the model can re-bind it | subject_token: "the woman in the red coat" |
| 3 | Scene content | Describes the new scene the subject is placed into | scene: "walking through a Tokyo alley at night, neon reflections on wet pavement" |
| 4 | Motion / camera direction | Specifies how the scene moves | camera: "slow dolly-in, eye-level, 24fps cinematic" |
| 5 | Audio cue | Sets sound design (or silence) | audio: "rain ambience, distant city hum, footsteps" |
The skeleton is portable because every commercial Ref2V API exposes these five slots, even under different names. Seedance wraps the anchor in @tag, Veo uses image_prompt, Kling uses elements[] with role keys, Wan uses start_image/end_image. The five components remain.
Keep the subject token description short — 2–6 words. Long natural-language subject descriptions re-introduce the ambiguity Ref2V removes.
4. Model-by-model template adaptations
Here is how the same Ref2V intent translates across the documented APIs, using a single-subject example: a woman in a red coat walking through a Tokyo alley at night.
| Component | Seedance 2.5 | Veo 3.1 | Kling 3.0 | Wan 3.0 |
|---|---|---|---|---|
| Source anchor | @subject (image attached) |
image_prompt: <image> |
elements[0].image |
start_image: <image> |
| Subject token | the woman in the red coat |
described in prompt field |
elements[0].description |
described in prompt field |
| Scene content | inline in prompt | prompt field |
prompt field |
prompt field |
| Motion / camera | inline after // camera: |
inline in prompt |
inline in prompt |
end_image + camera in prompt |
| Audio cue | inline // audio: |
audio_prompt field |
audio field |
not exposed; placeholder |
Three observations:
- Seedance and Veo are the most expressive. Seedance’s
@tagsystem is the most ergonomic for multi-reference; Veo’simage_promptis the cleanest single-subject surface. - Kling’s
elementsarray is the most explicit role-binding mechanism — strongest for multi-subject reference where subject bleed is the dominant risk. - Wan’s start/end frame is the most restrictive — you must specify the first and last frame, which is a strength for motion-driven reference but a limitation for free-form Ref2V.
This table is the spine of the playbook. When a template does not call out a specific model, the universal skeleton applies; adapt per this table.
5. The 12 prompt templates
Each template is a fenced text block. Model-specific adaptations follow in Section 6.
5.1 Single-subject reference (3 templates)
Template 1 — Walking portrait (single subject, urban)
# Ref2V — Single-subject walking portrait
# Identity-locked B-roll
[anchor]
reference_image: ./woman_red_coat.png
[subject]
subject_token: "the woman in the red coat"
preserve: face, hair length, coat color, coat cut
[scene]
location: Tokyo, Shinjuku alley, night
environment: wet pavement, neon (blue and magenta), light rain, steam from vents
time: 23:40
[motion]
camera: slow dolly-in, eye-level, 35mm feel
subject_motion: walks toward camera, slight head turn at the end
duration: 6s
fps: 24
[audio]
ambience: rain on pavement, distant traffic
foley: footsteps, coat rustle
Template 2 — Product showcase (single subject, studio)
# Ref2V — Single-subject product showcase
# E-commerce, brand lock
[anchor]
reference_image: ./ceramic_vase.png
[subject]
subject_token: "the matte-black ceramic vase"
preserve: silhouette, glaze finish, base footprint
[scene]
location: studio, seamless off-white backdrop
environment: soft key light from camera-left, subtle rim from right
[motion]
camera: slow 180-degree turntable, slight push-in at midpoint
subject_motion: static, camera orbits
duration: 5s
fps: 30
[audio]
ambience: silent room tone
Template 3 — Character dialogue beat (single subject, performance)
# Ref2V — Single-subject dialogue beat
# Narrative, emotional close-up
[anchor]
reference_image: ./character_portrait.png
[subject]
subject_token: "the man with the grey beard"
preserve: face, beard shape, eye color
[scene]
location: dimly lit diner booth
environment: warm tungsten key, cool window fill
time: 02:15
[motion]
camera: locked medium close-up, very subtle handheld drift
subject_motion: slight head tilt down, then up; eyes meet lens at the end
duration: 4s
fps: 24
[audio]
ambience: diner kitchen clatter, soft jazz underscore
foley: shallow breathing
5.2 Multi-subject reference (3 templates)
Template 4 — Two-character conversation
# Ref2V — Two-character conversation
# Dialogue scene with identity preservation
[anchors]
reference_image_a: ./character_A.png
reference_image_b: ./character_B.png
[subjects]
subject_token_a: "the woman with the short black hair"
subject_token_b: "the man in the grey cardigan"
preserve: face, hair, primary garment color per subject
[scene]
location: park bench, autumn afternoon
environment: golden-hour side key, leaf fall, clear and cool
time: 17:45
[motion]
camera: slow push-in from wide to medium two-shot
subject_motion: A gestures with right hand while speaking; B nods, looks down, then back up
duration: 7s
fps: 24
[audio]
ambience: distant children, wind in leaves
foley: rustling jacket, bench creak
Template 5 — Subject + prop handoff
# Ref2V — Subject + prop handoff
# Brand + actor, unboxing narrative
[anchors]
reference_image_subject: ./presenter.png
reference_image_prop: ./product_box.png
[subjects]
subject_token_subject: "the presenter in the navy blazer"
subject_token_prop: "the teal product box with the silver logo"
preserve: subject face and blazer cut; prop dimensions, logo placement, color saturation
[scene]
location: clean studio set
environment: high-key soft lighting, gradient grey backdrop
[motion]
camera: static medium shot, focus pull subject → prop at handoff
subject_motion: subject lifts box from table with both hands, holds chest-high, looks down; prop held still, slight reflection shift
duration: 5s
fps: 30
[audio]
ambience: silent
foley: cardboard slide on table
music: corporate underscore, fades in
Template 6 — Ensemble cast (three subjects)
# Ref2V — Ensemble cast
# Group shot, foreground/background separation
[anchors]
reference_image_a: ./cast_A.png
reference_image_b: ./cast_B.png
reference_image_c: ./cast_C.png
[subjects]
subject_token_a: "the older man with the round glasses"
subject_token_b: "the young woman with the red scarf"
subject_token_c: "the child in the denim jacket"
preserve: face and signature accessory per subject
[scene]
location: subway platform, midday
environment: fluorescent overhead, daylight wall behind
time: 12:10
[motion]
camera: static wide shot
subject_motion: a stands left, hands in pockets, slight sway; b stands center, scrolls phone; c stands right, looks up tunnel, points
duration: 6s
fps: 24
[audio]
ambience: train rumble, station announcement (indistinct)
foley: footsteps, paper rustle
5.3 Style-only reference (3 templates)
Template 7 — Cinematic color grade
# Ref2V — Style-only: cinematic color grade
# Lock a teal-and-orange look
[anchor]
reference_image: ./reference_grade.png
[subject]
preserve: none (style-only)
[style]
inherit:
palette: teal shadows, orange highlights, crushed blacks
contrast: high
grain: medium 35mm
halation: subtle warm bloom on highlights
do_not_inherit:
composition: ignore layout of reference
subjects: do not copy figures from reference
[scene]
scene_subject: a courier on a motorcycle riding through city traffic at dusk
[motion]
camera: tracking shot, slightly low angle
duration: 5s
fps: 24
[audio]
ambience: traffic hum
Template 8 — Painterly style transfer
# Ref2V — Style-only: painterly style
# Art-direction lock
[anchor]
reference_image: ./painting_reference.png
[subject]
preserve: scene composition, but render as if painted
[style]
inherit:
medium: oil on canvas
brushwork: visible, loose in background, tighter on focal subject
palette: muted earth tones with one accent (the painting's accent)
texture: canvas grain visible throughout
do_not_inherit:
subjects: do not copy figures; re-render new subjects in the same style
[scene]
scene_subject: a baker arranging bread in a window display, morning light
[motion]
camera: locked medium shot
duration: 4s
fps: 24
[audio]
ambience: city street outside
Template 9 — Brand visual identity
# Ref2V — Style-only: brand visual identity
# Campaign footage matching brand deck
[anchor]
reference_image: ./brand_moodboard.png
[subject]
preserve: brand palette, typography cue, motion cadence
[style]
inherit:
palette: brand primary (deep indigo) + accent (warm gold), off-white negative
composition: generous negative space, subject lower-third
motion: slow, deliberate, 1.2x slower than real time
text_treatment: leave room for headline overlay (lower-third safe area)
do_not_inherit:
subjects: do not copy figures from moodboard
props: do not copy specific products
[scene]
scene_subject: a small business owner unlocking their shop door at sunrise
[motion]
camera: static wide, then slow push-in during final second
duration: 6s
fps: 30
[audio]
ambience: morning birds
music: brand sonic logo sting at the end (placeholder)
5.4 Motion-driven reference (3 templates)
Template 10 — Start/end frame motion
# Ref2V — Motion-driven: start/end frame
# Wan-style first/last frame anchoring
[anchors]
start_image: ./frame_start.png # standing still
end_image: ./frame_end.png # mid-stride, arm raised
[subject]
subject_token: "the woman in the olive jacket"
preserve: face, jacket, bag strap
[scene]
location: city crosswalk, daytime, dry
environment: bright overcast, soft shadows
time: 14:30
[motion]
camera: locked eye-level medium shot
subject_motion: interpolates from standing still to mid-stride with raised arm
duration: 3s
fps: 30
[audio]
ambience: crosswalk signal, distant traffic
foley: footsteps
Template 11 — Camera motion only
# Ref2V — Motion-driven: camera-only
# Parallax reveal, locked subject
[anchor]
reference_image: ./interior_room.png
[subject]
subject_token: "the room interior (no specific human subject)"
preserve: composition, furniture placement, lighting
[scene]
location: loft apartment, late afternoon
environment: window light from camera-right
time: 18:10
[motion]
camera: slow dolly from doorway toward window, slight arc
subject_motion: none (subject is the room itself)
duration: 7s
fps: 24
[audio]
ambience: city hum outside window
music: ambient pad, low volume
Template 12 — Action arc with subject motion
# Ref2V — Motion-driven: action arc
# Dynamic beat with identity lock
[anchor]
reference_image: ./dancer.png
[subject]
subject_token: "the dancer in the white shirt"
preserve: face, shirt, hairstyle
[scene]
location: empty studio, single spotlight from above
environment: black backdrop, dust in the spotlight cone
[motion]
camera: handheld, orbits the subject in a half-circle during the move
subject_motion: leaps from a crouch, spins once mid-air, lands in a wide stance
duration: 4s
fps: 60
[audio]
ambience: silent studio
foley: shoe squeak, landing thud
music: none (post)
6. Model-by-model adaptations of the 12 templates
Below is Template 1 (walking portrait) translated into each model’s API surface. Shared scene text (referenced as <scene> below): “the woman in the red coat walks toward camera through a Tokyo alley at night. Wet pavement reflects blue and magenta neon. Light rain. Steam from vents.” Shared camera: “slow dolly-in, eye-level, 35mm cinematic.”
Seedance 2.5
@subject: ./woman_red_coat.png
<scene>
camera: slow dolly-in, eye-level, 35mm cinematic feel.
// audio: rain on pavement, distant traffic, footsteps, coat fabric
@subject binds the image to the nearest noun phrase; // comments are ignored; add @subject2, @subject3 for multi-subject.
Veo 3.1
{
"image_prompt": "./woman_red_coat.png",
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic",
"audio_prompt": "rain on pavement, distant traffic, footsteps, coat fabric"
}
image_prompt is the anchor; the rest goes into text fields. Keep scene descriptions under ~120 words.
Kling 3.0
{
"elements": [
{"image": "./woman_red_coat.png", "description": "the woman in the red coat", "role": "primary_subject"}
],
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic",
"audio": "rain on pavement, distant traffic, footsteps, coat fabric"
}
elements is the most explicit role-binding surface. For style-only, set role: "style_reference" and leave description empty.
Wan 3.0
{
"start_image": "./woman_red_coat.png",
"end_image": "./woman_red_coat_frame2.png",
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic"
}
Wan requires both start_image and end_image for motion-driven reference. Audio is not exposed natively; add in post.
Adaptation matrix for all 12 templates
| Template | Seedance 2.5 | Veo 3.1 | Kling 3.0 | Wan 3.0 | Notes |
|---|---|---|---|---|---|
| 1. Walking portrait | @subject |
image_prompt |
elements[0] |
start_image+end_image |
Wan needs hand-drawn end frame |
| 2. Product showcase | @subject+turntable |
image_prompt |
elements[0] |
start_image |
All four handle |
| 3. Dialogue beat | @subject |
image_prompt |
elements[0] |
weak fit | Wan weak on subtle expression |
| 4. Two-character | @subject+@subject2 |
two image_prompt |
elements[0..1] |
not supported | Kling wins |
| 5. Subject+prop | @subject+@prop |
two image_prompt |
elements[0..1] |
weak fit | Seedance/Kling strong |
| 6. Ensemble cast | @subject x3 |
not practical | elements[0..2] |
not supported | Kling only one-shot |
| 7. Cinematic grade | @style |
image_prompt+cue |
elements[0] style |
start_image+cue |
All four handle style |
| 8. Painterly style | @style |
image_prompt+medium |
elements[0] style |
start_image |
All four handle |
| 9. Brand identity | @style+brand cue |
image_prompt+cue |
elements[0] style |
start_image+cue |
All four handle |
| 10. Start/end frame | not native | not native | not native | native | Wan wins |
| 11. Camera-only | @scene+camera |
image_prompt+camera |
elements[0]+camera |
weak fit | Veo/Seedance strong |
| 12. Action arc | @subject+motion |
image_prompt+motion |
elements[0]+motion |
both frames | All four, with caveats |
7. Sister articles and further reading
- Ref2V companion article — Seedance Ref2V walkthrough.
- Seedance 2.5 prompts — Seedance 2.5 model guide.
- Seedance multi-reference image prompts — 50-asset companion.
- Image-to-video character lock prompts — adjacent I2V.
- Multimodal reference video prompts — multimodal context.
- Best AI video prompts 2026 masterclass — cross-model overview.
External sources:
- Ref2V paper (arXiv:2508.02458).
- Seedance 50 references guide.
- Veo documentation.
- AI filmmaking prompting guide.
8. Verifying H3 compatibility
A reproducible method to score H3 against the five-component skeleton:
Establish what H3 accepts. Inspect the H3 API for a reference image field (
image_prompt,reference_image,init_image,@tag, orelements[]), a subject token mechanism, style-only support, and start/end frame support. If you cannot confirm all four, treat H3 as partially specified.Start with Template 7 (cinematic grade). Style-only, cheapest to validate. A model that fails style-only fails everything.
Run Template 1 (walking portrait). The canonical single-subject Ref2V test.
Run Template 4 (two-character conversation). The first test where subject bleed shows up. If H3 swaps faces or merges features, that is a clear failure signal.
Score each output on five dimensions (0 = ignored, 1 = partial, 2 = full):
| Dimension | What to check | |———–|—————| | Anchor fidelity | Does the subject look like the reference? | | Style fidelity | Does the style match (if style-only)? | | Scene compliance | Does the scene match the written scene? | | Motion compliance | Does the motion match the camera/subject direction? | | Audio compliance | Does the audio match (if supported)? |
8+ out of 10 means H3 is compatible with the universal skeleton; below 8 means H3 needs its own adaptations.
- Document failures. A failed component (e.g., audio) is a feature gap, not a verdict on H3 as a whole. Note the failure and use templates that do not require it.
FAQ
What is a reference-to-video prompt template?
A reference-to-video (Ref2V) prompt template pairs a reference image with structured text so a video model produces output that inherits visual properties — identity, style, brand, or motion — from the reference. The canonical structure has five components: anchor, subject token, scene, motion, audio. See arXiv:2508.02458.
How is Ref2V different from I2V?
I2V animates a single source image and predicts plausible motion. Ref2V treats the input as a reference — a constraint on identity, style, or composition — paired with text that defines the new scene. Use Ref2V when you need identity lock or brand fidelity; I2V when you just want a still image to move.
Which model is best for multi-subject reference?
Kling 3.0 via its elements array with per-subject role keys. Seedance 2.5 supports it through @subject tags and is the most ergonomic inline. Veo 3.1 requires separate image_prompt fields and compositing. Wan 3.0 does not support multi-subject in one call.
Can I use these templates for H3 or S2V/V2V?
Designed against documented APIs and not verified against H3; field names may differ. The five-component skeleton should translate, but follow Section 8 to score H3. These target Ref2V specifically — S2V and V2V are related but the syntax does not carry over.
Is this article appropriate for H3?
Yes, as a starting point. We have not tested H3. Treat the templates as hypotheses, run Section 8, and adjust syntax to match H3’s actual API surface. We will update once we have independent H3 data.
How long should a Ref2V prompt be?
Keep scene descriptions under ~120 words and subject tokens to 2–6 words. Long descriptions reintroduce the ambiguity Ref2V removes.
Conclusion
Reference-to-video prompting in 2026 is a five-component skeleton that travels across Seedance, Veo, Kling, Wan, and — once verified — likely H3 and other forthcoming models. Treat the 12 templates as starting points, score your outputs against Section 8, and adapt syntax to whatever API surface you are working against.
For Seedance deep dives, see the Seedance 2.5 prompts guide and the multi-reference 50-asset companion. For multimodal framing, see multimodal reference video prompts.
Reviewed by the videosprompt.org editorial team · October 2026
Limitations of this article
The 12 templates are validated against the documented APIs of Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0. H3 is referenced in the title and slug for keyword discoverability, but we have not independently run H3 against these templates and lack first-party data on H3’s reference-image API surface. Readers applying them to H3 should treat the templates as hypotheses and follow Section 8. Results on H3 may differ from the documented models. We will revise once independent H3 data is available.
Share Article