Seedance 参考图生视频(Ref2V)提示词完全指南(附完整示例)
TL;DR
- Ref2V(Reference-to-Video,参考图生视频) 是 arXiv:2508.02458 提出的以主体为核心的生成框架,旨在仅凭一张参考图加一段目标文本提示,让特定人物、产品或风格在每一帧中保持一致。
- 与传统的 图生视频(I2V) 主要依据源图驱动运动不同,Ref2V 优先保障 主体身份一致性 —— 这是当前 AI 视频领域最难解决的问题。
- Seedance 2.5 支持最多 50 个多模态参考素材(30 张图片、10 段视频、10 段音频),通过
@tag语法调用,是 2026 年最具生产力的 Ref2V 流水线之一。 - 本指南包含 12+ 即拷即用提示词模板,覆盖播报讲解、人物舞蹈、产品主视觉和分镜故事板等场景。
- 提供完整的 防漂移提示词清单,帮助你在多片段作品中锁定发型、瞳色、服饰层次、配饰与光线。
- 文末附有内部链接:Seedance 2.5 提示词指南、Kling 3.0 人物一致性、2K 视频模型提示词模板 以及 2026 年最佳 AI 视频提示词大师课。
1. 什么是 Ref2V —— 以及它为何重要
如果你已经生成过不止几条 AI 视频,几乎一定遇到过同一个问题:刚让模型生成一个四秒片段,人物的脸就开始变形。第 3 帧到第 7 帧之间头发颜色发生了变化。原本的蓝色夹克到了结尾变成了绿色。前景里的产品在第 90 帧漂移成了另一个品牌。
这就是 主体漂移(subject drift),也是当前 AI 视频生成的首要难题。
Ref2V —— 即 Reference-to-Video(参考图生视频) —— 是论文 *Ref2V: When Your Image Is the Reference, Which Your Text Is The Target? A Joint Generation Framework for Reference-based Video Generation*(arXiv:2508.02458)正式提出的一种生成范式。其核心思想简洁而强大:
你提供 一张 主体参考图(人脸、角色、产品或风格画面)和 一段文本提示描述希望发生什么。模型生成的视频中,参考图中的主体在每一帧保持一致,而文本提示负责控制动作、镜头和场景。
这与传统的 I2V 范式有本质区别。在 I2V 中,源图设定 起始帧,模型据此推断动作。由于模型并未绑定特定身份,而是绑定在某个起始帧上,主体通常会发生漂移。Ref2V 颠覆了这一思路:参考图被视为 持续的身份锚点,文本提示被视为 需要达到的目标状态。每一帧都同时受两者约束。
对于做播报讲解、品牌锁定产品视频、多场景分镜或带重复角色的 MV 的人来说,Ref2V 决定了这条流水线是否真正可用。
2. Ref2V 与传统 I2V 的差异
Ref2V 与 I2V 的区别是 2026 年最常被混淆的知识点之一,而搞清楚这一点会彻底改变你写提示词的方式。
| 维度 | 传统 I2V | Ref2V |
|---|---|---|
| 输入 | 1 张源图 + 文本提示 | 1 张参考图(身份锚点)+ 文本提示 |
| 目标 | 按文本提示让源帧动起来 | 在每一帧保留参考主体,再执行文本 |
| 主体漂移容忍度 | 高 —— 面部、服装、颜色可自由变化 | 低 —— 主体在整段片段中锁定 |
| 最佳场景 | 通用运动、环境氛围镜头、风格化片段 | 人物驱动内容、品牌内容、分镜故事板 |
| 常见失败模式 | 动作常偏离文本提示 | 提示与参考冲突时身份会模糊 |
| 提示词重心 | 动作动词、镜头调度、场景氛围 | 主体锚点、防漂移线索、标签纪律 |
实操含义:当你写 Ref2V 提示词时,实际上是在 同时写两份契约 —— 一份是身份契约(哪些东西必须保持),一份是动作契约(需要发生什么)。I2V 提示词只需要承担后者的职责。
如果你正从 I2V 迁移到 Ref2V,最常见的错误是仍然只按驱动运动的方式写提示词。你会在前 2 秒内就看到漂移,因为你没告诉模型 哪些东西不能变。
3. Seedance 2.5 的参考系统
Seedance 2.5 是 ByteDance 推出的生产级视频模型,拥有市面上最慷慨的参考系统之一。根据 Seedance 2.5 发布报道 和 Seedance 官方博客,单次生成可摄取最多 50 个多模态参考素材:
- 30 张图片参考 —— 用于人物、产品、环境与风格画面。
- 10 段视频参考 —— 用于动作迁移、前序片段或实拍风格画面。
- 10 段音频参考 —— 用于声音克隆、背景音乐和音效分轨。
这些参考通过提示词中的 @tag 语法调用。为每个素材分配一个标签(例如 @char1、@product_a、@bg_studio、@vo_main),然后在叙事描述中引用它。
@tag 语法
@char1: a 28-year-old woman with shoulder-length auburn hair,
green eyes, light freckles, wearing a cream knit
sweater and small gold hoop earrings.
@product_a: a matte-black ceramic coffee mug with a single
centered gold handle, photographed on a white
seamless backdrop.
[Scene] — 2-second intro shot. @char1 holds @product_a at chest
height, rotating it 15 degrees so the gold handle
catches the key light. Soft top-down studio lighting,
shallow depth of field, 60mm equivalent.
有三点需要注意:
- 顶部的
@tag块是描述性的,不是叙事性的。 它规定 必须保持什么。方括号[Scene]部分才是动作发生的地方。 - 参考是锚点,不是演员。 模型知道
@product_a在第 0 帧和第 90 帧是同一个哑光黑色马克杯。 - 标签本身没有顺序。 如果你在场景中互换
@char1和@char2,模型就会把身份互换。标签纪律是 Ref2V 卫生中最大的问题。
想深入了解 Seedance 完整的提示词语法(风格 token、镜头控制、动作动词),请参阅 Seedance 2.5 提示词指南。
4. 三种参考模式
Ref2V 可以清晰地归为三种生产模式。大多数真实作品会在同一内容中同时用到全部三种。
4.1 单主体模式
一张参考图、一个身份、一个角色或产品贯穿整段片段。这是播报讲解、产品主视觉、单一角色模式的典型场景,也是防漂移提示词容错空间最小的模式。
4.2 多主体合奏模式
两张或更多参考图,每个代表不同主体,组合在同一个场景里。这是双人访谈、家庭场景、多产品阵列模式的典型场景。模型既要锁定每个主体的身份,又要让彼此保持区别 —— 这是一项严格更难的任务。
4.3 风格迁移模式
参考图贡献的是 视觉风格 而非具体主体。常见例子:把电影静帧、画作或品牌色调的视觉风格迁移到新场景上。参考图是 风格锚点,而非 身份锚点,主体与动作由文本提示提供。
5. 提示词模板
以下示例全部使用 Seedance 2.5 的 @tag 语法编写,每个都是即拷即用模板 —— 把方括号中的 [xxx] 替换为你自己的主体、场景与动作即可。
5.1 播报讲解模板
模板 1:直视镜头、中性主光
@host: A 30-year-old male presenter with short black hair,
clean-shaven, brown eyes, wearing a navy blue
crew-neck t-shirt, light skin, no visible jewelry.
[Scene] — Single-camera talking-head shot, locked-off
tripod framing from mid-chest up. @host looks directly
into camera and speaks with a calm, measured tone.
Subtle head movements and natural blink. Soft 3-point
studio lighting: warm key at 45 degrees camera left,
cool fill at 30 percent intensity, hair light from
behind. Static background: muted olive wall. Shot on
35mm equivalent, f/2.8, shallow DOF.
模板 2:手势强调、较亮主光
@host: A 45-year-old female executive with chin-length
silver hair, dark brown eyes, small pearl earrings,
wearing a charcoal pinstripe blazer over a white
silk blouse. Medium-brown skin tone.
[Scene] — Medium shot, locked tripod. @host stands
behind a clear acrylic podium against a deep teal
backdrop. Soft top-front key at 5600K, gentle rim
from camera right. @host speaks while gesturing with
both hands in front of her torso to emphasize points.
Minimal body movement otherwise.
模板 3:边走边说、外景
@host: A 25-year-old non-binary creator with a buzzcut
dyed lavender, hazel eyes, a small nose ring, and
a tan canvas jacket. Medium skin tone.
[Scene] — Handheld follow shot, 50mm equivalent,
f/4. @host walks toward camera on a city sidewalk
lined with brick storefronts, overcast daylight,
slight motion on the camera. @host speaks to camera
while walking. Autumn leaves on the ground. Jacket
remains visible throughout; wind does not flip the
collar.
5.2 角色舞蹈 / MV 模板
模板 4:独舞、固定编排
@dancer: A 22-year-old woman with jet-black hair in a
high ponytail, dark brown eyes, winged eyeliner,
wearing a red sequin halter top and black
high-waisted trousers. Light skin tone.
[Scene] — Wide shot, static camera. @dancer performs
a contemporary dance solo in the center of a dark
industrial loft, single overhead spotlight casting
a hard circular pool of light. Background is a
black void. Sequins catch the spotlight as she
moves. No camera movement. 24fps filmic motion blur
on hands only.
模板 5:双人舞、对称构图
@dancer_a: A 30-year-old man with a shaved head,
dark brown eyes, gold chain necklace,
wearing a sleeveless white tank top and
black joggers. Medium-dark skin tone.
@dancer_b: A 28-year-old woman with long box braids,
brown eyes, no jewelry, wearing a matching
sleeveless white tank top and black joggers.
Medium-dark skin tone.
[Scene] — Static wide shot, perfectly symmetrical
composition. @dancer_a and @dancer_b perform a
synchronized hip-hop duet in a sunlit warehouse with
dusty afternoon sun streaming through tall windows
behind them. Camera does not move. Outfits must
remain identical in shade and fabric for every frame.
模板 6:群舞编排、合奏
@lead: A 24-year-old woman with waist-length curly
red hair, green eyes, freckles, wearing a
black leather jacket over a white crop top.
Pale skin tone.
@ensemble: Five background dancers in matching black
tank tops and gray sweatpants, alternating
medium skin tones, no facial jewelry.
[Scene] — Wide crane-down shot starting high above
a neon-soaked Tokyo backstreet at night. Camera
slowly descends as @lead sings at center frame
while @ensemble forms a V-shape behind her.
Magenta and cyan neon signs in background. Rain
falls gently. @lead's leather jacket remains
glossy black throughout.
5.3 产品主视觉与品牌一致性模板
模板 7:单品、360° 旋转展示
@product: A 12oz matte-black ceramic coffee mug with
a single centered gold-colored handle,
branded with a small white wordmark logo
on the front. Photographed on a white
seamless backdrop.
[Scene] — Macro product hero shot. @product rotates
a full 360 degrees on a turntable over 4 seconds,
centered in frame. Soft diffused top-front key light,
white reflection beneath. Shallow DOF, 100mm
equivalent. The gold handle and the white wordmark
logo remain identical throughout — color, position,
and size do not shift between frames.
模板 8:产品使用、模特手
@product: A 16oz stainless steel water bottle in
brushed silver finish, laser-etched logo
on the lower third.
@hand: A 30-year-old woman's right hand, neutral
manicure with a clear coat, no rings.
[Scene] — Medium close-up. @hand reaches into frame
from the right and picks up @product from a
marble countertop. @hand lifts @product to mouth
level and tilts it as if drinking. Soft window
light from camera left, marble is white with light
gray veining. @hand releases @product back onto
the marble. Logo on @product remains visible and
in the same position on the bottle throughout.
模板 9:多产品对比
@product_a: A wireless over-ear headphone in matte
black, silver hinges, brand wordmark
on the ear cup.
@product_b: An identical model in pearl white, gold
hinges, brand wordmark on the ear cup.
[Scene] — Locked-off product shot on a warm beige
seamless backdrop. @product_a sits at frame left,
@product_b at frame right, both at the same height.
Soft top-front key light, slight rim from camera
right. A hand enters frame from the bottom and
rotates @product_a 20 degrees to display the
hinge, then exits. @product_b does not move.
Both wordmarks remain legible throughout.
5.4 同一角色跨场景分镜模板
模板 10:单一角色、三个场景
@hero: A 35-year-old man with short sandy-blonde hair,
blue eyes, a trimmed beard, and a small scar
above his left eyebrow. Wearing a faded brown
leather jacket over a gray henley shirt.
Light skin tone.
[Scene 1] — Exterior, golden hour. @hero stands in
a wheat field, wind blows the wheat but the
leather jacket does not flap dramatically. Wide
shot, locked tripod, 35mm equivalent.
[Scene 2] — Interior, low-key lighting. @hero sits
at a diner counter, fluorescent light overhead. Medium
close-up, shallow focus on @hero's face, the
darker background blurs. Same jacket, same scar,
same beard.
[Scene 3] — Exterior, night. @hero walks down a
rain-slicked alley, neon signs reflecting on
wet pavement. Handheld follow shot from behind.
Same jacket, now slightly wet but still the
same brown color, no pattern shift.
模板 11:双角色、保持连续性
@protagonist: A 40-year-old woman with shoulder-
length straight black hair, dark brown
eyes, red lipstick, wearing a fitted
navy blue trench coat. Medium skin
tone.
@antagonist: A 50-year-old man with silver hair
slicked back, gray eyes, clean-shaven,
wearing a charcoal cashmere overcoat.
Pale skin tone.
[Scene] — Exterior, night, rain. @protagonist and
@antagonist face each other on an empty city
street, 4 feet apart. Streetlamp overhead between
them creates a hard rim light on both faces.
Static wide shot, 24mm equivalent, deep focus.
Both coats remain dry on the upper shoulders,
water beads on the surface. @protagonist's red
lipstick stays the same shade in every frame.
模板 12:主角跨昼夜弧线
@hero: A 28-year-old non-binary person with a
short fade haircut dyed copper-orange,
brown eyes, wearing an oversized cream
wool sweater and round tortoiseshell
glasses. Medium skin tone.
[Scene 1 — Dawn] — Exterior rooftop. Pink and
orange sky. @hero sits on the ledge looking
out. Wide shot. Wool sweater is cream.
[Scene 2 — Midday] — Same rooftop. Bright sun
directly overhead. @hero stands, looking down
at the city. Medium-full. Same cream sweater,
no yellowing from the sun.
[Scene 3 — Dusk] — Same rooftop. Orange and
purple sky. @hero walks toward camera. Same
cream sweater, glasses still on, hair still
copper-orange even as the ambient light
shifts toward warm. Close-up, 85mm equivalent.
关于更多跨场景连续性写法,请参阅 Kling 3.0 人物一致性指南 —— 相同的防漂移技巧在不同模型家族中同样适用。
6. 防漂移提示词清单
漂移发生的根源是模型需要 猜测 哪些东西应当保持不变。你的工作是消除这种猜测。你每写一条防漂移线索,就是给模型一个可用的约束。
6.1 发型
说明长度(齐肩、平头、齐腰)、颜色(赭色、淡紫、沙金色)和发质(直发、卷发、螺卷、辫发)。如果发型样式很重要 —— 刘海、铲青、偏分 —— 务必写明。长发的提示中需说明风是否会影响发型。
6.2 瞳色
永远要写。瞳色是当前视频模型中漂移最快的属性之一。在 @tag 定义里写一次,再通过描述眼神方向(向上看、向左瞥一眼)隐性强化。
6.3 服装层次
列出参考图中所有可见的层次:外层、中间层、内层,以及任何可见的内衣边缘。如果角色在参考图里敞开夹克,提示词中必须说明它保持敞开还是合上。不要假设模型会跨帧维持某种层叠状态。
6.4 配饰
耳环、项链、眼镜、手表、发夹、戒指。所有可见的配饰都应被命名一次。配饰漂移很微妙 —— 第 60 帧手表消失、第 90 帧耳环不见 —— 但观众潜意识里最先察觉的恰恰是这些细节。
6.5 光线连续性
同一场景内的光线应当在 @tag 块或场景设置里明确。多场景之间保持光线连续更难 —— 但你可以锁定一个 方向(始终是相机右侧的轮廓光)和 色温(始终 5600K)。这能防止模型在场景推进时漂移到不同的情绪。
6.6 图案与质感保持
条纹、Logo、亮片、刺绣、织物纹路。这些都属于高漂移元素,因为模型倾向于 规则化 图案。在场景描述中要再次强调(”亮片捕捉到聚光灯”、”Logo 在同一位置保持清晰”)。
6.7 5 秒法则
如果场景描述里第 90 帧时主体没有变化,就再写一次。“同样的夹克,同样的颜色,没有图案漂移。” 这在 token 上只有很小的冗余成本,却能显著降低漂移。
关于锁定身份提示词写法的更深入讨论,请参阅 2026 年最佳 AI 视频提示词大师课。
7. 常见错误
7.1 主体描述过于模糊
“一个女人”或”一个穿夹克的人”等于没有给模型任何可锁定信息。每少写一个描述,模型就多出一个自由度,而每一帧它都会填上不同的答案。
7.2 互换 @tag 分配
如果 @character1 在场景 1 中引用、@character2 在场景 2 中引用,但两者的描述恰好互换,模型就会把身份互相对调。标签分配是身份契约的一部分。
7.3 风格线索互相冲突
在同一段提示词中同时写”真实感光照”和”动漫风格头发”不是简单的画面描述,而是反复触发的漂移源头。选择单一视觉风格并保持一致。对于风格迁移模式,应使用专门的 @style 参考图,而非在线索中混搭。
7.4 描述模型无法完成的运动
Ref2V 不会凭空赋予超越基础视频模型本身的运动能力。如果你要求一段 12 秒、180° 弧线的跟踪镜头,其实是在要求底层模型做它可能根本做不到的事。运动指令必须与底层模型文档承诺的能力对齐。
7.5 沿用 I2V 的提示词习惯
只写动作和场景(”镜头平移,她微笑”),而不指明哪些东西不能变,这是从 I2V 迁移过来最常见的错误。如果场景里没有写明身份,模型就无从锁定。
7.6 忽略音频参考
Seedance 2.5 的音频参考(最多 10 个)属于同一套 @tag 系统。如果你想在多场景作品中保持人声连续,只需要在场景中持续引用 @vo_main。音频漂移和视觉漂移一样真实存在。
8. 常见问题
Ref2V 和 I2V 在实操中的区别是什么?
在 I2V 中,源图就是第一帧,模型可以自由演化主体。在 Ref2V 中,参考图是贯穿每一帧的持久身份锚点,文本则描述围绕它发生的事。简单来说:Ref2V 让面部、产品和服装保持稳定;I2V 则允许它们变形。
Ref2V 能用多张参考图吗?
可以。Seedance 2.5 单次生成支持最多 30 张图片参考,每张都对应自己的 @tag。多主体合奏(双人角色、多产品)是 Ref2V 的标准模式之一。
视频和音频参考各能用多少?
视频参考最多 10 个,音频参考最多 10 个,加上 30 张图片参考 —— 单次生成总计 50 个多模态参考素材。
提示词最长能写多少?
Seedance 2.5 的提示词可以轻松容纳数百个 token,包括 @tag 定义和场景描述。5 秒法则(在第 90 帧再次描述保持不变的内容)完全在预算范围内。
Ref2V 能保证零漂移吗?
不能。相比 I2V,Ref2V 大幅降低 漂移,是当前最强的身份一致性方案,但 2026 年没有任何生产级模型能保证长片段零漂移。本指南中的防漂移清单能帮你最大程度接近零漂移。
可以把 Ref2V 和风格迁移结合起来吗?
可以 —— 这就是第三种参考模式。用 @style 参考图承载视觉风格,用 @subject 参考图承载身份,再用文本描述动作。避免在线索中把风格和身份混在一起写。
Ref2V 跟其他模型里的人物一致性工具是同一回事吗?
底层 目标 一致 —— 让主体在帧间保持稳定。实现机制 不同。Ref2V 是 arXiv:2508.02458 论文中正式化的训练范式;而 Kling、Runway、Sora 等模型中的人物一致性是通过不同的架构选择实现的。提示词的原则可以互通。
在哪里能找到更多提示词模板?
我们的 2K 视频模型提示词模板 和 Seedance 2.5 提示词指南 是最直接相关的延伸阅读。
9. 结语
Ref2V 是一次范式跃迁,把 AI 视频从尝鲜循环改造成真正的生产流水线。”一张参考图 + 目标文本”的格式在 arXiv:2508.02458 中正式化,并在 Seedance 2.5 的 50 参考素材系统中达到生产规模,让创作者能够把一张脸、一个产品或一种风格锁定到故事需要的任意多片段上。
本指南中的模板只是起点,让它们真正可用于生产的是防漂移技巧。把每一个 @tag 视为一份身份契约,把每一段场景描述同时视为动作契约 和 对”什么不能变”的再次声明。这种纪律正是把零散片段与成系列作品区分开的关键。
延伸阅读:Seedance 2.5 提示词指南、Kling 3.0 人物一致性详解、2K 视频模型提示词模板 以及 2026 年最佳 AI 视频提示词大师课。
参考资料
- Ref2V: When Your Image Is the Reference, Which Your Text Is The Target? (arXiv:2508.02458)
- Seedance 官方博客
- Seedance 2.5 发布报道 —— 50 参考素材
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章