VideosPrompt VideosPrompt

Reference-to-Video Templates: A Cross-Model Playbook for Seedance, Veo, Kling, Wan, and Beyond

Author: VideosPrompt Date: 2026-10-07 14:11:43
Reference-to-Video Templates: A Cross-Model Playbook for Seedance, Veo, Kling, Wan, and Beyond

TL;DR

  • Reference-to-video (Ref2V) uses one or more still images as the visual anchor for a generated video clip, rather than relying on text alone.
  • The three core Ref2V techniques are single-subject reference, multi-subject reference, and style-only reference, each with distinct scaffolds and failure modes.
  • This article ships 12 copy-ready templates adapted for Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0.
  • A universal Ref2V skeleton — anchor, subject token, scene, motion/camera, audio — travels across all four model APIs.
  • MiniMax H3 is referenced but unverified. We have not run H3 against these templates; the playbooks are validated against the four documented APIs with a checklist for H3.
  • Style-only references are the cheapest to validate and the best starting point for any new model including H3.
  • A verification protocol at the end lets you score any model, including H3, against the same five-component skeleton.

1. What “reference-to-video” means in 2026

Reference-to-video (Ref2V) conditions a video model on one or more input images so the output inherits specific visual properties — a face, a product, a lighting style, a brand palette — that text alone cannot reliably describe. The Ref2V paper (arXiv:2508.02458) formalized the subject-token handoff pattern that all major commercial APIs now expose under different names.

Three core techniques sit inside Ref2V:

Technique Primary use case Inputs Failure mode
Single-subject reference Lock one identity (face, product, character) across a clip 1 reference image + scene description Identity drift when scene prompt dominates the reference
Multi-subject reference Composite two or more characters/objects into one scene 2–N reference images + role tags Subject bleed and pose confusion when tokens are not disambiguated
Style-only reference Inherit lighting, palette, or artistic treatment 1 reference image + style-only prompt Style overpowers subject; subjects get re-stylized rather than re-rendered

The distinction matters because each technique has a different prompt skeleton, cost profile, and risk of failure. A practical Ref2V playbook handles all three without collapsing them.


2. What we verified, what we couldn’t

Verified against documented APIs. The 12 templates are adapted against publicly documented surfaces of four models:

For prompt engineering context, see the mstudio.ai filmmaking prompting guide.

MiniMax H3 is referenced but unverified. The keyword for this piece names H3, so we include it in the title for discoverability. We have not run H3 against these templates and do not know whether H3 uses @tag-style syntax, image_prompt, elements[], or another mechanism. Apply these to H3 as a hypothesis and use Section 8 to score actual output.

We are not going to fabricate H3-specific field names or claim benchmark numbers we did not run.


3. The Ref2V prompt skeleton

Every Ref2V prompt decomposes into five components. We use this skeleton across all 12 templates because it survives the translation between APIs better than model-native syntax does.

# Component Purpose Universal example
1 Source reference anchor Tells the model which image(s) to consult reference_image: ./subject.png
2 Subject token description Names what the anchor represents so the model can re-bind it subject_token: "the woman in the red coat"
3 Scene content Describes the new scene the subject is placed into scene: "walking through a Tokyo alley at night, neon reflections on wet pavement"
4 Motion / camera direction Specifies how the scene moves camera: "slow dolly-in, eye-level, 24fps cinematic"
5 Audio cue Sets sound design (or silence) audio: "rain ambience, distant city hum, footsteps"

The skeleton is portable because every commercial Ref2V API exposes these five slots, even under different names. Seedance wraps the anchor in @tag, Veo uses image_prompt, Kling uses elements[] with role keys, Wan uses start_image/end_image. The five components remain.

Keep the subject token description short — 2–6 words. Long natural-language subject descriptions re-introduce the ambiguity Ref2V removes.


4. Model-by-model template adaptations

Here is how the same Ref2V intent translates across the documented APIs, using a single-subject example: a woman in a red coat walking through a Tokyo alley at night.

Component Seedance 2.5 Veo 3.1 Kling 3.0 Wan 3.0
Source anchor @subject (image attached) image_prompt: <image> elements[0].image start_image: <image>
Subject token the woman in the red coat described in prompt field elements[0].description described in prompt field
Scene content inline in prompt prompt field prompt field prompt field
Motion / camera inline after // camera: inline in prompt inline in prompt end_image + camera in prompt
Audio cue inline // audio: audio_prompt field audio field not exposed; placeholder

Three observations:

  1. Seedance and Veo are the most expressive. Seedance’s @tag system is the most ergonomic for multi-reference; Veo’s image_prompt is the cleanest single-subject surface.
  2. Kling’s elements array is the most explicit role-binding mechanism — strongest for multi-subject reference where subject bleed is the dominant risk.
  3. Wan’s start/end frame is the most restrictive — you must specify the first and last frame, which is a strength for motion-driven reference but a limitation for free-form Ref2V.

This table is the spine of the playbook. When a template does not call out a specific model, the universal skeleton applies; adapt per this table.


5. The 12 prompt templates

Each template is a fenced text block. Model-specific adaptations follow in Section 6.

5.1 Single-subject reference (3 templates)

Template 1 — Walking portrait (single subject, urban)

# Ref2V — Single-subject walking portrait
# Identity-locked B-roll

[anchor]
reference_image: ./woman_red_coat.png

[subject]
subject_token: "the woman in the red coat"
preserve: face, hair length, coat color, coat cut

[scene]
location: Tokyo, Shinjuku alley, night
environment: wet pavement, neon (blue and magenta), light rain, steam from vents
time: 23:40

[motion]
camera: slow dolly-in, eye-level, 35mm feel
subject_motion: walks toward camera, slight head turn at the end
duration: 6s
fps: 24

[audio]
ambience: rain on pavement, distant traffic
foley: footsteps, coat rustle

Template 2 — Product showcase (single subject, studio)

# Ref2V — Single-subject product showcase
# E-commerce, brand lock

[anchor]
reference_image: ./ceramic_vase.png

[subject]
subject_token: "the matte-black ceramic vase"
preserve: silhouette, glaze finish, base footprint

[scene]
location: studio, seamless off-white backdrop
environment: soft key light from camera-left, subtle rim from right

[motion]
camera: slow 180-degree turntable, slight push-in at midpoint
subject_motion: static, camera orbits
duration: 5s
fps: 30

[audio]
ambience: silent room tone

Template 3 — Character dialogue beat (single subject, performance)

# Ref2V — Single-subject dialogue beat
# Narrative, emotional close-up

[anchor]
reference_image: ./character_portrait.png

[subject]
subject_token: "the man with the grey beard"
preserve: face, beard shape, eye color

[scene]
location: dimly lit diner booth
environment: warm tungsten key, cool window fill
time: 02:15

[motion]
camera: locked medium close-up, very subtle handheld drift
subject_motion: slight head tilt down, then up; eyes meet lens at the end
duration: 4s
fps: 24

[audio]
ambience: diner kitchen clatter, soft jazz underscore
foley: shallow breathing

5.2 Multi-subject reference (3 templates)

Template 4 — Two-character conversation

# Ref2V — Two-character conversation
# Dialogue scene with identity preservation

[anchors]
reference_image_a: ./character_A.png
reference_image_b: ./character_B.png

[subjects]
subject_token_a: "the woman with the short black hair"
subject_token_b: "the man in the grey cardigan"
preserve: face, hair, primary garment color per subject

[scene]
location: park bench, autumn afternoon
environment: golden-hour side key, leaf fall, clear and cool
time: 17:45

[motion]
camera: slow push-in from wide to medium two-shot
subject_motion: A gestures with right hand while speaking; B nods, looks down, then back up
duration: 7s
fps: 24

[audio]
ambience: distant children, wind in leaves
foley: rustling jacket, bench creak

Template 5 — Subject + prop handoff

# Ref2V — Subject + prop handoff
# Brand + actor, unboxing narrative

[anchors]
reference_image_subject: ./presenter.png
reference_image_prop: ./product_box.png

[subjects]
subject_token_subject: "the presenter in the navy blazer"
subject_token_prop: "the teal product box with the silver logo"
preserve: subject face and blazer cut; prop dimensions, logo placement, color saturation

[scene]
location: clean studio set
environment: high-key soft lighting, gradient grey backdrop

[motion]
camera: static medium shot, focus pull subject → prop at handoff
subject_motion: subject lifts box from table with both hands, holds chest-high, looks down; prop held still, slight reflection shift
duration: 5s
fps: 30

[audio]
ambience: silent
foley: cardboard slide on table
music: corporate underscore, fades in

Template 6 — Ensemble cast (three subjects)

# Ref2V — Ensemble cast
# Group shot, foreground/background separation

[anchors]
reference_image_a: ./cast_A.png
reference_image_b: ./cast_B.png
reference_image_c: ./cast_C.png

[subjects]
subject_token_a: "the older man with the round glasses"
subject_token_b: "the young woman with the red scarf"
subject_token_c: "the child in the denim jacket"
preserve: face and signature accessory per subject

[scene]
location: subway platform, midday
environment: fluorescent overhead, daylight wall behind
time: 12:10

[motion]
camera: static wide shot
subject_motion: a stands left, hands in pockets, slight sway; b stands center, scrolls phone; c stands right, looks up tunnel, points
duration: 6s
fps: 24

[audio]
ambience: train rumble, station announcement (indistinct)
foley: footsteps, paper rustle

5.3 Style-only reference (3 templates)

Template 7 — Cinematic color grade

# Ref2V — Style-only: cinematic color grade
# Lock a teal-and-orange look

[anchor]
reference_image: ./reference_grade.png

[subject]
preserve: none (style-only)

[style]
inherit:
  palette: teal shadows, orange highlights, crushed blacks
  contrast: high
  grain: medium 35mm
  halation: subtle warm bloom on highlights
do_not_inherit:
  composition: ignore layout of reference
  subjects: do not copy figures from reference

[scene]
scene_subject: a courier on a motorcycle riding through city traffic at dusk

[motion]
camera: tracking shot, slightly low angle
duration: 5s
fps: 24

[audio]
ambience: traffic hum

Template 8 — Painterly style transfer

# Ref2V — Style-only: painterly style
# Art-direction lock

[anchor]
reference_image: ./painting_reference.png

[subject]
preserve: scene composition, but render as if painted

[style]
inherit:
  medium: oil on canvas
  brushwork: visible, loose in background, tighter on focal subject
  palette: muted earth tones with one accent (the painting's accent)
  texture: canvas grain visible throughout
do_not_inherit:
  subjects: do not copy figures; re-render new subjects in the same style

[scene]
scene_subject: a baker arranging bread in a window display, morning light

[motion]
camera: locked medium shot
duration: 4s
fps: 24

[audio]
ambience: city street outside

Template 9 — Brand visual identity

# Ref2V — Style-only: brand visual identity
# Campaign footage matching brand deck

[anchor]
reference_image: ./brand_moodboard.png

[subject]
preserve: brand palette, typography cue, motion cadence

[style]
inherit:
  palette: brand primary (deep indigo) + accent (warm gold), off-white negative
  composition: generous negative space, subject lower-third
  motion: slow, deliberate, 1.2x slower than real time
  text_treatment: leave room for headline overlay (lower-third safe area)
do_not_inherit:
  subjects: do not copy figures from moodboard
  props: do not copy specific products

[scene]
scene_subject: a small business owner unlocking their shop door at sunrise

[motion]
camera: static wide, then slow push-in during final second
duration: 6s
fps: 30

[audio]
ambience: morning birds
music: brand sonic logo sting at the end (placeholder)

5.4 Motion-driven reference (3 templates)

Template 10 — Start/end frame motion

# Ref2V — Motion-driven: start/end frame
# Wan-style first/last frame anchoring

[anchors]
start_image: ./frame_start.png   # standing still
end_image: ./frame_end.png       # mid-stride, arm raised

[subject]
subject_token: "the woman in the olive jacket"
preserve: face, jacket, bag strap

[scene]
location: city crosswalk, daytime, dry
environment: bright overcast, soft shadows
time: 14:30

[motion]
camera: locked eye-level medium shot
subject_motion: interpolates from standing still to mid-stride with raised arm
duration: 3s
fps: 30

[audio]
ambience: crosswalk signal, distant traffic
foley: footsteps

Template 11 — Camera motion only

# Ref2V — Motion-driven: camera-only
# Parallax reveal, locked subject

[anchor]
reference_image: ./interior_room.png

[subject]
subject_token: "the room interior (no specific human subject)"
preserve: composition, furniture placement, lighting

[scene]
location: loft apartment, late afternoon
environment: window light from camera-right
time: 18:10

[motion]
camera: slow dolly from doorway toward window, slight arc
subject_motion: none (subject is the room itself)
duration: 7s
fps: 24

[audio]
ambience: city hum outside window
music: ambient pad, low volume

Template 12 — Action arc with subject motion

# Ref2V — Motion-driven: action arc
# Dynamic beat with identity lock

[anchor]
reference_image: ./dancer.png

[subject]
subject_token: "the dancer in the white shirt"
preserve: face, shirt, hairstyle

[scene]
location: empty studio, single spotlight from above
environment: black backdrop, dust in the spotlight cone

[motion]
camera: handheld, orbits the subject in a half-circle during the move
subject_motion: leaps from a crouch, spins once mid-air, lands in a wide stance
duration: 4s
fps: 60

[audio]
ambience: silent studio
foley: shoe squeak, landing thud
music: none (post)

6. Model-by-model adaptations of the 12 templates

Below is Template 1 (walking portrait) translated into each model’s API surface. Shared scene text (referenced as <scene> below): “the woman in the red coat walks toward camera through a Tokyo alley at night. Wet pavement reflects blue and magenta neon. Light rain. Steam from vents.” Shared camera: “slow dolly-in, eye-level, 35mm cinematic.”

Seedance 2.5

@subject: ./woman_red_coat.png
<scene>
camera: slow dolly-in, eye-level, 35mm cinematic feel.
// audio: rain on pavement, distant traffic, footsteps, coat fabric

@subject binds the image to the nearest noun phrase; // comments are ignored; add @subject2, @subject3 for multi-subject.

Veo 3.1

{
  "image_prompt": "./woman_red_coat.png",
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic",
  "audio_prompt": "rain on pavement, distant traffic, footsteps, coat fabric"
}

image_prompt is the anchor; the rest goes into text fields. Keep scene descriptions under ~120 words.

Kling 3.0

{
  "elements": [
    {"image": "./woman_red_coat.png", "description": "the woman in the red coat", "role": "primary_subject"}
  ],
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic",
  "audio": "rain on pavement, distant traffic, footsteps, coat fabric"
}

elements is the most explicit role-binding surface. For style-only, set role: "style_reference" and leave description empty.

Wan 3.0

{
  "start_image": "./woman_red_coat.png",
  "end_image": "./woman_red_coat_frame2.png",
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic"
}

Wan requires both start_image and end_image for motion-driven reference. Audio is not exposed natively; add in post.

Adaptation matrix for all 12 templates

Template Seedance 2.5 Veo 3.1 Kling 3.0 Wan 3.0 Notes
1. Walking portrait @subject image_prompt elements[0] start_image+end_image Wan needs hand-drawn end frame
2. Product showcase @subject+turntable image_prompt elements[0] start_image All four handle
3. Dialogue beat @subject image_prompt elements[0] weak fit Wan weak on subtle expression
4. Two-character @subject+@subject2 two image_prompt elements[0..1] not supported Kling wins
5. Subject+prop @subject+@prop two image_prompt elements[0..1] weak fit Seedance/Kling strong
6. Ensemble cast @subject x3 not practical elements[0..2] not supported Kling only one-shot
7. Cinematic grade @style image_prompt+cue elements[0] style start_image+cue All four handle style
8. Painterly style @style image_prompt+medium elements[0] style start_image All four handle
9. Brand identity @style+brand cue image_prompt+cue elements[0] style start_image+cue All four handle
10. Start/end frame not native not native not native native Wan wins
11. Camera-only @scene+camera image_prompt+camera elements[0]+camera weak fit Veo/Seedance strong
12. Action arc @subject+motion image_prompt+motion elements[0]+motion both frames All four, with caveats

7. Sister articles and further reading

External sources:


8. Verifying H3 compatibility

A reproducible method to score H3 against the five-component skeleton:

  1. Establish what H3 accepts. Inspect the H3 API for a reference image field (image_prompt, reference_image, init_image, @tag, or elements[]), a subject token mechanism, style-only support, and start/end frame support. If you cannot confirm all four, treat H3 as partially specified.

  2. Start with Template 7 (cinematic grade). Style-only, cheapest to validate. A model that fails style-only fails everything.

  3. Run Template 1 (walking portrait). The canonical single-subject Ref2V test.

  4. Run Template 4 (two-character conversation). The first test where subject bleed shows up. If H3 swaps faces or merges features, that is a clear failure signal.

  5. Score each output on five dimensions (0 = ignored, 1 = partial, 2 = full):

| Dimension | What to check | |———–|—————| | Anchor fidelity | Does the subject look like the reference? | | Style fidelity | Does the style match (if style-only)? | | Scene compliance | Does the scene match the written scene? | | Motion compliance | Does the motion match the camera/subject direction? | | Audio compliance | Does the audio match (if supported)? |

8+ out of 10 means H3 is compatible with the universal skeleton; below 8 means H3 needs its own adaptations.

  1. Document failures. A failed component (e.g., audio) is a feature gap, not a verdict on H3 as a whole. Note the failure and use templates that do not require it.

FAQ

What is a reference-to-video prompt template?

A reference-to-video (Ref2V) prompt template pairs a reference image with structured text so a video model produces output that inherits visual properties — identity, style, brand, or motion — from the reference. The canonical structure has five components: anchor, subject token, scene, motion, audio. See arXiv:2508.02458.

How is Ref2V different from I2V?

I2V animates a single source image and predicts plausible motion. Ref2V treats the input as a reference — a constraint on identity, style, or composition — paired with text that defines the new scene. Use Ref2V when you need identity lock or brand fidelity; I2V when you just want a still image to move.

Which model is best for multi-subject reference?

Kling 3.0 via its elements array with per-subject role keys. Seedance 2.5 supports it through @subject tags and is the most ergonomic inline. Veo 3.1 requires separate image_prompt fields and compositing. Wan 3.0 does not support multi-subject in one call.

Can I use these templates for H3 or S2V/V2V?

Designed against documented APIs and not verified against H3; field names may differ. The five-component skeleton should translate, but follow Section 8 to score H3. These target Ref2V specifically — S2V and V2V are related but the syntax does not carry over.

Is this article appropriate for H3?

Yes, as a starting point. We have not tested H3. Treat the templates as hypotheses, run Section 8, and adjust syntax to match H3’s actual API surface. We will update once we have independent H3 data.

How long should a Ref2V prompt be?

Keep scene descriptions under ~120 words and subject tokens to 2–6 words. Long descriptions reintroduce the ambiguity Ref2V removes.


Conclusion

Reference-to-video prompting in 2026 is a five-component skeleton that travels across Seedance, Veo, Kling, Wan, and — once verified — likely H3 and other forthcoming models. Treat the 12 templates as starting points, score your outputs against Section 8, and adapt syntax to whatever API surface you are working against.

For Seedance deep dives, see the Seedance 2.5 prompts guide and the multi-reference 50-asset companion. For multimodal framing, see multimodal reference video prompts.


Reviewed by the videosprompt.org editorial team · October 2026


Limitations of this article

The 12 templates are validated against the documented APIs of Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0. H3 is referenced in the title and slug for keyword discoverability, but we have not independently run H3 against these templates and lack first-party data on H3’s reference-image API surface. Readers applying them to H3 should treat the templates as hypotheses and follow Section 8. Results on H3 may differ from the documented models. We will revise once independent H3 data is available.

Share Article

Related Articles

Recommended Reading

Ready to Get Started?

Experience our product immediately and explore more possibilities.