Reference-to-Video (Ref2V) Prompts for Seedance: A Practical Guide with Full Examples
TL;DR
- Ref2V (Reference-to-Video) is the subject-driven video generation framework introduced in arXiv:2508.02458, designed to keep a specific character, product, or style locked across every frame from a single reference image plus a target text prompt.
- Unlike vanilla image-to-video (I2V), which primarily animates motion from a source image, Ref2V prioritizes subject identity preservation — the single hardest problem in modern AI video.
- Seedance 2.5 supports up to 50 multimodal reference assets (30 images, 10 video clips, 10 audio clips) accessible through an
@tagsyntax, making it one of the most production-ready Ref2V pipelines available in 2026. - This guide includes 12+ copy-paste prompt templates across talking-head, character dance, product hero, and storyboard use cases.
- We include a full anti-drift prompting checklist so you can lock hair, eye color, clothing layers, accessories, and lighting across multi-clip productions.
- Internal links to our Seedance 2.5 prompts guide, Kling 3.0 character consistency, 2K video prompt templates, and best AI video prompts 2026 masterclass are included throughout.
1. What Ref2V is — and why it matters
If you have generated more than a handful of AI videos, you have almost certainly run into the same problem: the moment you ask the model for a four-second clip, the character’s face morphs. The hair color shifts between frames three and seven. A blue jacket becomes a green jacket by the end. The product in the foreground drifts to a different brand at frame 90.
This is subject drift, and it is the #1 problem in AI video generation today.
Ref2V — short for Reference-to-Video — is a generation paradigm formalized in the paper Ref2V: When Your Image Is the Reference, Which Your Text Is The Target? A Joint Generation Framework for Reference-based Video Generation (arXiv:2508.02458). The central idea is simple but powerful:
You supply one reference image of a subject (a face, a character, a product, a style frame) and a text prompt describing what should happen. The model generates a video in which the subject from the reference image is preserved across every frame, while the text prompt controls motion, action, camera, and environment.
This is materially different from the older I2V paradigm. In I2V, the source image sets the first frame and the model extrapolates motion. The subject can — and usually does — drift because the model is not bound to any specific identity; it is bound to a single starting frame. Ref2V inverts this. The reference image is treated as a persistent identity anchor, and the text prompt is treated as the target state to reach. Every frame is conditioned on both.
For anyone building talking-head explainers, brand-locked product videos, multi-scene storyboards, or music videos with a recurring character, Ref2V is the difference between a usable pipeline and a constant manual-fix loop.
2. How Ref2V differs from vanilla I2V
The distinction between Ref2V and I2V is one of the most common points of confusion in 2026, and getting it right changes how you write prompts.
| Dimension | Vanilla I2V | Ref2V |
|---|---|---|
| Input | 1 source image + text prompt | 1 reference image (identity anchor) + text prompt |
| Goal | Animate the source frame per the text | Preserve subject from reference across all frames, then act on the text |
| Subject drift tolerance | High — face, clothing, and color can morph freely | Low — subject is locked across the entire clip |
| Best for | Generic motion, ambient shots, stylistic clips | Character-driven content, branded content, storyboards |
| Failure mode | Action often diverges from prompt | Identity can soften if prompt contradicts reference |
| Prompt writing focus | Motion verbs, camera moves, scene mood | Subject anchors, anti-drift cues, tag discipline |
The practical implication: when you write a Ref2V prompt, you are writing two things at once — an identity contract (what must stay) and an action contract (what must happen). I2V prompts only need to deliver the second.
If you are migrating from I2V to Ref2V, the most common mistake is to keep writing as if you were driving motion only. You will see drift within the first 2 seconds because you have not told the model what must not change.
3. Seedance 2.5’s reference system
Seedance 2.5 is ByteDance’s production-grade video model with one of the most generous reference systems on the market. According to the Seedance 2.5 launch coverage and Seedance’s official blog, a single generation can ingest up to 50 multimodal reference assets:
- 30 image references — for subjects, products, environments, and style frames.
- 10 video references — for motion transfer, prior clips in a sequence, or live-action style frames.
- 10 audio references — for voice clone, background music beds, and SFX stems.
These references are addressed through a simple @tag syntax inside the prompt. You assign each asset a tag (e.g., @char1, @product_a, @bg_studio, @vo_main) and then refer to it in your narrative description.
The @tag syntax
@char1: a 28-year-old woman with shoulder-length auburn hair,
green eyes, light freckles, wearing a cream knit
sweater and small gold hoop earrings.
@product_a: a matte-black ceramic coffee mug with a single
centered gold handle, photographed on a white
seamless backdrop.
[Scene] — 2-second intro shot. @char1 holds @product_a at chest
height, rotating it 15 degrees so the gold handle
catches the key light. Soft top-down studio lighting,
shallow depth of field, 60mm equivalent.
Three things to notice:
- The
@tagblock at the top is descriptive, not narrative. It establishes what must persist. The bracketed[Scene]section is where motion happens. - References are anchors, not actors. A sticker knows
@product_ais the same matte-black mug at frame 0 and frame 90. - Tags carry no inherent order. If you swap
@char1and@char2in the scene, the model swaps the identities. Tag discipline is the single biggest Ref2V hygiene issue.
For a deeper dive into Seedance’s broader prompt grammar (style tokens, camera controls, motion verbs), see our Seedance 2.5 prompts guide.
4. The three reference modes
Ref2V collapses cleanly into three production modes. Most real productions use all three within a single piece of content.
4.1 Single-subject mode
One reference image, one identity, one character or product across the entire clip. This is the talking-head, product-hero, and solo-character mode. It is also where anti-drift prompting has the smallest margin for error.
4.2 Multi-subject ensemble mode
Two or more reference images, each representing a distinct subject, combined in a single scene. This is the duo-interview, family-scene, and multi-product-lineup mode. The model must keep each subject’s identity locked and keep them distinct from each other — a strictly harder task.
4.3 Style transfer mode
The reference image contributes a visual style rather than a specific subject. Common examples: transferring the look of a film still, a painting, or a brand’s color palette onto a new scene. The reference is a style anchor, not an identity anchor. The prompt text provides the subject and the action.
5. Prompt Templates
All examples below are written for Seedance 2.5’s @tag syntax. Each is a copy-paste-ready template — swap the bracketed [xxx] text for your own subject, scene, and motion.
5.1 Talking-head explainer templates
Template 1: Direct address, neutral key light
@host: A 30-year-old male presenter with short black hair,
clean-shaven, brown eyes, wearing a navy blue
crew-neck t-shirt, light skin, no visible jewelry.
[Scene] — Single-camera talking-head shot, locked-off
tripod framing from mid-chest up. @host looks directly
into camera and speaks with a calm, measured tone.
Subtle head movements and natural blink. Soft 3-point
studio lighting: warm key at 45 degrees camera left,
cool fill at 30 percent intensity, hair light from
behind. Static background: muted olive wall. Shot on
35mm equivalent, f/2.8, shallow DOF.
Template 2: Hand-gesture emphasis, brighter key
@host: A 45-year-old female executive with chin-length
silver hair, dark brown eyes, small pearl earrings,
wearing a charcoal pinstripe blazer over a white
silk blouse. Medium-brown skin tone.
[Scene] — Medium shot, locked tripod. @host stands
behind a clear acrylic podium against a deep teal
backdrop. Soft top-front key at 5600K, gentle rim
from camera right. @host speaks while gesturing with
both hands in front of her torso to emphasize points.
Minimal body movement otherwise.
Template 3: Walk-and-talk, exterior
@host: A 25-year-old non-binary creator with a buzzcut
dyed lavender, hazel eyes, a small nose ring, and
a tan canvas jacket. Medium skin tone.
[Scene] — Handheld follow shot, 50mm equivalent,
f/4. @host walks toward camera on a city sidewalk
lined with brick storefronts, overcast daylight,
slight motion on the camera. @host speaks to camera
while walking. Autumn leaves on the ground. Jacket
remains visible throughout; wind does not flip the
collar.
5.2 Character dance / music video templates
Template 4: Solo dancer, locked choreography
@dancer: A 22-year-old woman with jet-black hair in a
high ponytail, dark brown eyes, winged eyeliner,
wearing a red sequin halter top and black
high-waisted trousers. Light skin tone.
[Scene] — Wide shot, static camera. @dancer performs
a contemporary dance solo in the center of a dark
industrial loft, single overhead spotlight casting
a hard circular pool of light. Background is a
black void. Sequins catch the spotlight as she
moves. No camera movement. 24fps filmic motion blur
on hands only.
Template 5: Duet, symmetrical framing
@dancer_a: A 30-year-old man with a shaved head,
dark brown eyes, gold chain necklace,
wearing a sleeveless white tank top and
black joggers. Medium-dark skin tone.
@dancer_b: A 28-year-old woman with long box braids,
brown eyes, no jewelry, wearing a matching
sleeveless white tank top and black joggers.
Medium-dark skin tone.
[Scene] — Static wide shot, perfectly symmetrical
composition. @dancer_a and @dancer_b perform a
synchronized hip-hop duet in a sunlit warehouse with
dusty afternoon sun streaming through tall windows
behind them. Camera does not move. Outfits must
remain identical in shade and fabric for every frame.
Template 6: Group choreography, ensemble
@lead: A 24-year-old woman with waist-length curly
red hair, green eyes, freckles, wearing a
black leather jacket over a white crop top.
Pale skin tone.
@ensemble: Five background dancers in matching black
tank tops and gray sweatpants, alternating
medium skin tones, no facial jewelry.
[Scene] — Wide crane-down shot starting high above
a neon-soaked Tokyo backstreet at night. Camera
slowly descends as @lead sings at center frame
while @ensemble forms a V-shape behind her.
Magenta and cyan neon signs in background. Rain
falls gently. @lead's leather jacket remains
glossy black throughout.
5.3 Product hero with consistent branding templates
Template 7: Single product, hero rotation
@product: A 12oz matte-black ceramic coffee mug with
a single centered gold-colored handle,
branded with a small white wordmark logo
on the front. Photographed on a white
seamless backdrop.
[Scene] — Macro product hero shot. @product rotates
a full 360 degrees on a turntable over 4 seconds,
centered in frame. Soft diffused top-front key light,
white reflection beneath. Shallow DOF, 100mm
equivalent. The gold handle and the white wordmark
logo remain identical throughout — color, position,
and size do not shift between frames.
Template 8: Product in use, hand model
@product: A 16oz stainless steel water bottle in
brushed silver finish, laser-etched logo
on the lower third.
@hand: A 30-year-old woman's right hand, neutral
manicure with a clear coat, no rings.
[Scene] — Medium close-up. @hand reaches into frame
from the right and picks up @product from a
marble countertop. @hand lifts @product to mouth
level and tilts it as if drinking. Soft window
light from camera left, marble is white with light
gray veining. @hand releases @product back onto
the marble. Logo on @product remains visible and
in the same position on the bottle throughout.
Template 9: Multi-product comparison
@product_a: A wireless over-ear headphone in matte
black, silver hinges, brand wordmark
on the ear cup.
@product_b: An identical model in pearl white, gold
hinges, brand wordmark on the ear cup.
[Scene] — Locked-off product shot on a warm beige
seamless backdrop. @product_a sits at frame left,
@product_b at frame right, both at the same height.
Soft top-front key light, slight rim from camera
right. A hand enters frame from the bottom and
rotates @product_a 20 degrees to display the
hinge, then exits. @product_b does not move.
Both wordmarks remain legible throughout.
5.4 Storyboard with same character across scenes templates
Template 10: Single character, three scenes
@hero: A 35-year-old man with short sandy-blonde hair,
blue eyes, a trimmed beard, and a small scar
above his left eyebrow. Wearing a faded brown
leather jacket over a gray henley shirt.
Light skin tone.
[Scene 1] — Exterior, golden hour. @hero stands in
a wheat field, wind blows the wheat but the
leather jacket does not flap dramatically. Wide
shot, locked tripod, 35mm equivalent.
[Scene 2] — Interior, low-key lighting. @hero sits
at a diner counter, fluorescent light overhead. Medium
close-up, shallow focus on @hero's face, the
darker background blurs. Same jacket, same scar,
same beard.
[Scene 3] — Exterior, night. @hero walks down a
rain-slicked alley, neon signs reflecting on
wet pavement. Handheld follow shot from behind.
Same jacket, now slightly wet but still the
same brown color, no pattern shift.
Template 11: Two-character scene with continuity
@protagonist: A 40-year-old woman with shoulder-
length straight black hair, dark brown
eyes, red lipstick, wearing a fitted
navy blue trench coat. Medium skin
tone.
@antagonist: A 50-year-old man with silver hair
slicked back, gray eyes, clean-shaven,
wearing a charcoal cashmere overcoat.
Pale skin tone.
[Scene] — Exterior, night, rain. @protagonist and
@antagonist face each other on an empty city
street, 4 feet apart. Streetlamp overhead between
them creates a hard rim light on both faces.
Static wide shot, 24mm equivalent, deep focus.
Both coats remain dry on the upper shoulders,
water beads on the surface. @protagonist's red
lipstick stays the same shade in every frame.
Template 12: Hero across time-of-day arc
@hero: A 28-year-old non-binary person with a
short fade haircut dyed copper-orange,
brown eyes, wearing an oversized cream
wool sweater and round tortoiseshell
glasses. Medium skin tone.
[Scene 1 — Dawn] — Exterior rooftop. Pink and
orange sky. @hero sits on the ledge looking
out. Wide shot. Wool sweater is cream.
[Scene 2 — Midday] — Same rooftop. Bright sun
directly overhead. @hero stands, looking down
at the city. Medium-full. Same cream sweater,
no yellowing from the sun.
[Scene 3 — Dusk] — Same rooftop. Orange and
purple sky. @hero walks toward camera. Same
cream sweater, glasses still on, hair still
copper-orange even as the ambient light
shifts toward warm. Close-up, 85mm equivalent.
For more cross-scene continuity patterns, see our guide to Kling 3.0 character consistency — the same anti-drift techniques apply across model families.
6. Anti-drift prompting checklist
Drift happens when the model has to guess what should stay the same. Your job is to remove the guess. Every anti-drift cue you write is a constraint the model can use.
6.1 Hair
Specify length (shoulder-length, buzzcut, waist-length), color (auburn, lavender, sandy-blonde), and texture (straight, curly, coiled, box braids). If the cut matters — bangs, undercut, side part — say so. For long hair, mention whether the wind affects it.
6.2 Eye color
Always. Eye color is one of the fastest drift points in modern video models. Specify it once in the @tag definition, then reinforce it implicitly by describing eye direction (looks up, glances left).
6.3 Clothing layers
List every visible layer in the reference image. Outerwear, mid-layer, base layer, and any visible undergarment edge. If the character has a jacket open in the reference, the prompt must specify whether it stays open or closes. Do not assume the model will hold a layer state across frames.
6.4 Accessories
Earrings, necklaces, glasses, watches, hair clips, rings. Every visible accessory should be named once. Accessory drift is subtle — a watch disappearing at frame 60, an earring vanishing at frame 90 — but it is one of the first things audiences pick up on subconsciously.
6.5 Lighting continuity
Within a scene, lighting should be specified in the @tag block or scene setup. Across scenes, lighting continuity is harder — but you can lock a direction (always rim from camera-right) and a color temperature (always 5600K). This prevents the model from drifting into a different mood as the scene progresses.
6.6 Pattern and texture preservation
Stripes, logos, sequins, embroidery, fabric weave. These are high-drift because the model tends to regularize patterns. Reinforce the pattern in the scene description (“sequins catch the spotlight,” “the wordmark remains legible in the same position”).
6.7 The 5-second rule
If the subject has not changed in the scene description at frame 90, write it again. “Same jacket, same color, no pattern shift.” This is a small redundancy cost in tokens but a large reduction in drift.
For a deeper treatment of prompt patterns that lock identity, see our best AI video prompts 2026 masterclass.
7. Common mistakes
7.1 Vague subject descriptions
“a woman” or “a person in a jacket” gives the model nothing to lock onto. Every descriptor you omit is a degree of freedom the model will fill — and it will fill each one differently across frames.
7.2 Swapping @tag assignments
If @character1 is referenced in scene 1 and @character2 in scene 2 but the descriptions swap, the model will swap the identities. Tag assignments are part of the identity contract.
7.3 Conflicting style cues
“Photorealistic lighting” in the same prompt as “anime-style hair” is not a single frame, it is a recurring drift trigger. Pick a single visual register and hold it. For style transfer mode, use a dedicated @style reference instead of mixing cues inline.
7.4 Describing motion the model cannot perform
Ref2V does not magically grant motion abilities beyond the base video model’s. If you ask for a 12-second tracking shot with a 180-degree arc, you are asking the base model to do something it may not be able to do. Match motion requests to what the underlying model is documented to produce.
7.5 Using I2V prompting habits
Writing only about motion and scene (“the camera pans, she smiles”) without naming what must not change is the #1 migration mistake from I2V. If you have not described the identity in the scene, the model has nothing to lock.
7.6 Ignoring audio references
Seedance 2.5’s audio references (up to 10) are part of the same @tag system. If you want voice continuity across a multi-scene piece, reference @vo_main once and let it persist across scenes. Audio drift is as real as visual drift.
8. FAQ
What is the difference between Ref2V and I2V in practice?
In I2V, the source image is the first frame and the model is free to evolve the subject. In Ref2V, the reference image is a persistent identity anchor across every frame, and the text describes what happens around it. Practically: Ref2V keeps faces, products, and outfits stable; I2V lets them morph.
Does Ref2V work with multiple reference images?
Yes. Seedance 2.5 supports up to 30 image references per generation, each with its own @tag. Multi-subject ensembles (two characters, multiple products) are a standard Ref2V mode.
How many video and audio references can I use?
Up to 10 video references and 10 audio references, in addition to the 30 image references — for a total of 50 multimodal reference assets per generation.
What is the maximum prompt length?
Seedance 2.5 prompts comfortably support several hundred tokens including @tag definitions and scene descriptions. The 5-second rule (“describe what stays the same again at frame 90”) is well within budget.
Does Ref2V guarantee zero drift?
No. Ref2V reduces drift dramatically compared to I2V and is the strongest current approach to subject consistency, but no production-grade model in 2026 guarantees zero drift on long clips. The anti-drift checklist in this guide is what gets you closest.
Can I combine Ref2V with style transfer?
Yes — this is the third reference mode. Use a @style reference image for the look and a @subject reference for the identity, then describe the action in text. Avoid mixing style cues inline with the identity description.
Is Ref2V the same as character consistency tools in other models?
The underlying goal is the same — keep a subject stable across frames. The mechanism differs. Ref2V is a training-and-inside paradigm formalized in arXiv:2508.02458; character consistency in other models (Kling, Runway, Sora) is achieved through different architectural choices. The prompting principles transfer.
Where can I find more prompt templates?
Our 2K video prompt templates and Seedance 2.5 prompts guide are the most directly relevant next reads.
9. Conclusion
Ref2V is the paradigm shift that turns AI video from a novelty loop into a production pipeline. The single reference image + target text format, formalized in arXiv:2508.02458 and realized at production scale in Seedance 2.5’s 50-reference-asset system, gives creators a way to lock a face, a product, or a style across as many clips as a story requires.
The templates in this guide are starting points. The anti-drift techniques are what make them production-ready. Treat every @tag as an identity contract and every scene description as both an action contract and a re-statement of what must not change. That discipline is what separates a one-off clip from a series.
For adjacent reading, see the Seedance 2.5 prompts guide, our Kling 3.0 character consistency breakdown, the 2K video model prompt templates, and the best AI video prompts 2026 masterclass.
Sources
- Ref2V: When Your Image Is the Reference, Which Your Text Is The Target? (arXiv:2508.02458)
- Seedance official blog
- Seedance 2.5 launch coverage — 50 reference assets
Reviewed by the videosprompt.org editorial team · October 2026
Share Article