VideosPrompt VideosPrompt

Seedance 参考图生视频(Ref2V)提示词完全指南(附完整示例)

作者: VideosPrompt 日期: 2026-10-07 14:11:54
Seedance 参考图生视频(Ref2V)提示词完全指南(附完整示例)

TL;DR

  • Ref2V(Reference-to-Video,参考图生视频) 是 arXiv:2508.02458 提出的以主体为核心的生成框架,旨在仅凭一张参考图加一段目标文本提示,让特定人物、产品或风格在每一帧中保持一致。
  • 与传统的 图生视频(I2V) 主要依据源图驱动运动不同,Ref2V 优先保障 主体身份一致性 —— 这是当前 AI 视频领域最难解决的问题。
  • Seedance 2.5 支持最多 50 个多模态参考素材(30 张图片、10 段视频、10 段音频),通过 @tag 语法调用,是 2026 年最具生产力的 Ref2V 流水线之一。
  • 本指南包含 12+ 即拷即用提示词模板,覆盖播报讲解、人物舞蹈、产品主视觉和分镜故事板等场景。
  • 提供完整的 防漂移提示词清单,帮助你在多片段作品中锁定发型、瞳色、服饰层次、配饰与光线。
  • 文末附有内部链接:Seedance 2.5 提示词指南、Kling 3.0 人物一致性、2K 视频模型提示词模板 以及 2026 年最佳 AI 视频提示词大师课。

1. 什么是 Ref2V —— 以及它为何重要

如果你已经生成过不止几条 AI 视频,几乎一定遇到过同一个问题:刚让模型生成一个四秒片段,人物的脸就开始变形。第 3 帧到第 7 帧之间头发颜色发生了变化。原本的蓝色夹克到了结尾变成了绿色。前景里的产品在第 90 帧漂移成了另一个品牌。

这就是 主体漂移(subject drift),也是当前 AI 视频生成的首要难题。

Ref2V —— 即 Reference-to-Video(参考图生视频) —— 是论文 *Ref2V: When Your Image Is the Reference, Which Your Text Is The Target? A Joint Generation Framework for Reference-based Video Generation*(arXiv:2508.02458)正式提出的一种生成范式。其核心思想简洁而强大:

你提供 一张 主体参考图(人脸、角色、产品或风格画面)和 一段文本提示描述希望发生什么。模型生成的视频中,参考图中的主体在每一帧保持一致,而文本提示负责控制动作、镜头和场景。

这与传统的 I2V 范式有本质区别。在 I2V 中,源图设定 起始帧,模型据此推断动作。由于模型并未绑定特定身份,而是绑定在某个起始帧上,主体通常会发生漂移。Ref2V 颠覆了这一思路:参考图被视为 持续的身份锚点,文本提示被视为 需要达到的目标状态。每一帧都同时受两者约束。

对于做播报讲解、品牌锁定产品视频、多场景分镜或带重复角色的 MV 的人来说,Ref2V 决定了这条流水线是否真正可用。


2. Ref2V 与传统 I2V 的差异

Ref2V 与 I2V 的区别是 2026 年最常被混淆的知识点之一,而搞清楚这一点会彻底改变你写提示词的方式。

维度 传统 I2V Ref2V
输入 1 张源图 + 文本提示 1 张参考图(身份锚点)+ 文本提示
目标 按文本提示让源帧动起来 在每一帧保留参考主体,再执行文本
主体漂移容忍度 高 —— 面部、服装、颜色可自由变化 低 —— 主体在整段片段中锁定
最佳场景 通用运动、环境氛围镜头、风格化片段 人物驱动内容、品牌内容、分镜故事板
常见失败模式 动作常偏离文本提示 提示与参考冲突时身份会模糊
提示词重心 动作动词、镜头调度、场景氛围 主体锚点、防漂移线索、标签纪律

实操含义:当你写 Ref2V 提示词时,实际上是在 同时写两份契约 —— 一份是身份契约(哪些东西必须保持),一份是动作契约(需要发生什么)。I2V 提示词只需要承担后者的职责。

如果你正从 I2V 迁移到 Ref2V,最常见的错误是仍然只按驱动运动的方式写提示词。你会在前 2 秒内就看到漂移,因为你没告诉模型 哪些东西不能变。


3. Seedance 2.5 的参考系统

Seedance 2.5 是 ByteDance 推出的生产级视频模型,拥有市面上最慷慨的参考系统之一。根据 Seedance 2.5 发布报道 和 Seedance 官方博客,单次生成可摄取最多 50 个多模态参考素材:

  • 30 张图片参考 —— 用于人物、产品、环境与风格画面。
  • 10 段视频参考 —— 用于动作迁移、前序片段或实拍风格画面。
  • 10 段音频参考 —— 用于声音克隆、背景音乐和音效分轨。

这些参考通过提示词中的 @tag 语法调用。为每个素材分配一个标签(例如 @char1、@product_a、@bg_studio、@vo_main),然后在叙事描述中引用它。

@tag 语法

@char1: a 28-year-old woman with shoulder-length auburn hair,
        green eyes, light freckles, wearing a cream knit
        sweater and small gold hoop earrings.

@product_a: a matte-black ceramic coffee mug with a single
            centered gold handle, photographed on a white
            seamless backdrop.

[Scene] — 2-second intro shot. @char1 holds @product_a at chest
         height, rotating it 15 degrees so the gold handle
         catches the key light. Soft top-down studio lighting,
         shallow depth of field, 60mm equivalent.

有三点需要注意:

  1. 顶部的 @tag 块是描述性的,不是叙事性的。 它规定 必须保持什么。方括号 [Scene] 部分才是动作发生的地方。
  2. 参考是锚点,不是演员。 模型知道 @product_a 在第 0 帧和第 90 帧是同一个哑光黑色马克杯。
  3. 标签本身没有顺序。 如果你在场景中互换 @char1 和 @char2,模型就会把身份互换。标签纪律是 Ref2V 卫生中最大的问题。

想深入了解 Seedance 完整的提示词语法(风格 token、镜头控制、动作动词),请参阅 Seedance 2.5 提示词指南。


4. 三种参考模式

Ref2V 可以清晰地归为三种生产模式。大多数真实作品会在同一内容中同时用到全部三种。

4.1 单主体模式

一张参考图、一个身份、一个角色或产品贯穿整段片段。这是播报讲解、产品主视觉、单一角色模式的典型场景,也是防漂移提示词容错空间最小的模式。

4.2 多主体合奏模式

两张或更多参考图,每个代表不同主体,组合在同一个场景里。这是双人访谈、家庭场景、多产品阵列模式的典型场景。模型既要锁定每个主体的身份,又要让彼此保持区别 —— 这是一项严格更难的任务。

4.3 风格迁移模式

参考图贡献的是 视觉风格 而非具体主体。常见例子:把电影静帧、画作或品牌色调的视觉风格迁移到新场景上。参考图是 风格锚点,而非 身份锚点,主体与动作由文本提示提供。


5. 提示词模板

以下示例全部使用 Seedance 2.5 的 @tag 语法编写,每个都是即拷即用模板 —— 把方括号中的 [xxx] 替换为你自己的主体、场景与动作即可。

5.1 播报讲解模板

模板 1:直视镜头、中性主光

@host: A 30-year-old male presenter with short black hair,
       clean-shaven, brown eyes, wearing a navy blue
       crew-neck t-shirt, light skin, no visible jewelry.

[Scene] — Single-camera talking-head shot, locked-off
tripod framing from mid-chest up. @host looks directly
into camera and speaks with a calm, measured tone.
Subtle head movements and natural blink. Soft 3-point
studio lighting: warm key at 45 degrees camera left,
cool fill at 30 percent intensity, hair light from
behind. Static background: muted olive wall. Shot on
35mm equivalent, f/2.8, shallow DOF.

模板 2:手势强调、较亮主光

@host: A 45-year-old female executive with chin-length
       silver hair, dark brown eyes, small pearl earrings,
       wearing a charcoal pinstripe blazer over a white
       silk blouse. Medium-brown skin tone.

[Scene] — Medium shot, locked tripod. @host stands
behind a clear acrylic podium against a deep teal
backdrop. Soft top-front key at 5600K, gentle rim
from camera right. @host speaks while gesturing with
both hands in front of her torso to emphasize points.
Minimal body movement otherwise.

模板 3:边走边说、外景

@host: A 25-year-old non-binary creator with a buzzcut
       dyed lavender, hazel eyes, a small nose ring, and
       a tan canvas jacket. Medium skin tone.

[Scene] — Handheld follow shot, 50mm equivalent,
f/4. @host walks toward camera on a city sidewalk
lined with brick storefronts, overcast daylight,
slight motion on the camera. @host speaks to camera
while walking. Autumn leaves on the ground. Jacket
remains visible throughout; wind does not flip the
collar.

5.2 角色舞蹈 / MV 模板

模板 4:独舞、固定编排

@dancer: A 22-year-old woman with jet-black hair in a
         high ponytail, dark brown eyes, winged eyeliner,
         wearing a red sequin halter top and black
         high-waisted trousers. Light skin tone.

[Scene] — Wide shot, static camera. @dancer performs
a contemporary dance solo in the center of a dark
industrial loft, single overhead spotlight casting
a hard circular pool of light. Background is a
black void. Sequins catch the spotlight as she
moves. No camera movement. 24fps filmic motion blur
on hands only.

模板 5:双人舞、对称构图

@dancer_a: A 30-year-old man with a shaved head,
           dark brown eyes, gold chain necklace,
           wearing a sleeveless white tank top and
           black joggers. Medium-dark skin tone.

@dancer_b: A 28-year-old woman with long box braids,
           brown eyes, no jewelry, wearing a matching
           sleeveless white tank top and black joggers.
           Medium-dark skin tone.

[Scene] — Static wide shot, perfectly symmetrical
composition. @dancer_a and @dancer_b perform a
synchronized hip-hop duet in a sunlit warehouse with
dusty afternoon sun streaming through tall windows
behind them. Camera does not move. Outfits must
remain identical in shade and fabric for every frame.

模板 6:群舞编排、合奏

@lead: A 24-year-old woman with waist-length curly
       red hair, green eyes, freckles, wearing a
       black leather jacket over a white crop top.
       Pale skin tone.

@ensemble: Five background dancers in matching black
           tank tops and gray sweatpants, alternating
           medium skin tones, no facial jewelry.

[Scene] — Wide crane-down shot starting high above
a neon-soaked Tokyo backstreet at night. Camera
slowly descends as @lead sings at center frame
while @ensemble forms a V-shape behind her.
Magenta and cyan neon signs in background. Rain
falls gently. @lead's leather jacket remains
glossy black throughout.

5.3 产品主视觉与品牌一致性模板

模板 7:单品、360° 旋转展示

@product: A 12oz matte-black ceramic coffee mug with
          a single centered gold-colored handle,
          branded with a small white wordmark logo
          on the front. Photographed on a white
          seamless backdrop.

[Scene] — Macro product hero shot. @product rotates
a full 360 degrees on a turntable over 4 seconds,
centered in frame. Soft diffused top-front key light,
white reflection beneath. Shallow DOF, 100mm
equivalent. The gold handle and the white wordmark
logo remain identical throughout — color, position,
and size do not shift between frames.

模板 8:产品使用、模特手

@product: A 16oz stainless steel water bottle in
          brushed silver finish, laser-etched logo
          on the lower third.

@hand: A 30-year-old woman's right hand, neutral
       manicure with a clear coat, no rings.

[Scene] — Medium close-up. @hand reaches into frame
from the right and picks up @product from a
marble countertop. @hand lifts @product to mouth
level and tilts it as if drinking. Soft window
light from camera left, marble is white with light
gray veining. @hand releases @product back onto
the marble. Logo on @product remains visible and
in the same position on the bottle throughout.

模板 9:多产品对比

@product_a: A wireless over-ear headphone in matte
            black, silver hinges, brand wordmark
            on the ear cup.

@product_b: An identical model in pearl white, gold
            hinges, brand wordmark on the ear cup.

[Scene] — Locked-off product shot on a warm beige
seamless backdrop. @product_a sits at frame left,
@product_b at frame right, both at the same height.
Soft top-front key light, slight rim from camera
right. A hand enters frame from the bottom and
rotates @product_a 20 degrees to display the
hinge, then exits. @product_b does not move.
Both wordmarks remain legible throughout.

5.4 同一角色跨场景分镜模板

模板 10:单一角色、三个场景

@hero: A 35-year-old man with short sandy-blonde hair,
      blue eyes, a trimmed beard, and a small scar
      above his left eyebrow. Wearing a faded brown
      leather jacket over a gray henley shirt.
      Light skin tone.

[Scene 1] — Exterior, golden hour. @hero stands in
a wheat field, wind blows the wheat but the
leather jacket does not flap dramatically. Wide
shot, locked tripod, 35mm equivalent.

[Scene 2] — Interior, low-key lighting. @hero sits
at a diner counter, fluorescent light overhead. Medium
close-up, shallow focus on @hero's face, the
darker  background blurs. Same jacket, same scar,
same beard.

[Scene 3] — Exterior, night. @hero walks down a
rain-slicked alley, neon signs reflecting on
wet pavement. Handheld follow shot from behind.
Same jacket, now slightly wet but still the
same brown color, no pattern shift.

模板 11:双角色、保持连续性

@protagonist: A 40-year-old woman with shoulder-
              length straight black hair, dark brown
              eyes, red lipstick, wearing a fitted
              navy blue trench coat. Medium skin
              tone.

@antagonist: A 50-year-old man with silver hair
             slicked back, gray eyes, clean-shaven,
             wearing a charcoal cashmere overcoat.
             Pale skin tone.

[Scene] — Exterior, night, rain. @protagonist and
@antagonist face each other on an empty city
street, 4 feet apart. Streetlamp overhead between
them creates a hard rim light on both faces.
Static wide shot, 24mm equivalent, deep focus.
Both coats remain dry on the upper shoulders,
water beads on the surface. @protagonist's red
lipstick stays the same shade in every frame.

模板 12:主角跨昼夜弧线

@hero: A 28-year-old non-binary person with a
      short fade haircut dyed copper-orange,
      brown eyes, wearing an oversized cream
      wool sweater and round tortoiseshell
      glasses. Medium skin tone.

[Scene 1 — Dawn] — Exterior rooftop. Pink and
orange sky. @hero sits on the ledge looking
out. Wide shot. Wool sweater is cream.

[Scene 2 — Midday] — Same rooftop. Bright sun
directly overhead. @hero stands, looking down
at the city. Medium-full. Same cream sweater,
no yellowing from the sun.

[Scene 3 — Dusk] — Same rooftop. Orange and
purple sky. @hero walks toward camera. Same
cream sweater, glasses still on, hair still
copper-orange even as the ambient light
shifts toward warm. Close-up, 85mm equivalent.

关于更多跨场景连续性写法,请参阅 Kling 3.0 人物一致性指南 —— 相同的防漂移技巧在不同模型家族中同样适用。


6. 防漂移提示词清单

漂移发生的根源是模型需要 猜测 哪些东西应当保持不变。你的工作是消除这种猜测。你每写一条防漂移线索,就是给模型一个可用的约束。

6.1 发型

说明长度(齐肩、平头、齐腰)、颜色(赭色、淡紫、沙金色)和发质(直发、卷发、螺卷、辫发)。如果发型样式很重要 —— 刘海、铲青、偏分 —— 务必写明。长发的提示中需说明风是否会影响发型。

6.2 瞳色

永远要写。瞳色是当前视频模型中漂移最快的属性之一。在 @tag 定义里写一次,再通过描述眼神方向(向上看、向左瞥一眼)隐性强化。

6.3 服装层次

列出参考图中所有可见的层次:外层、中间层、内层,以及任何可见的内衣边缘。如果角色在参考图里敞开夹克,提示词中必须说明它保持敞开还是合上。不要假设模型会跨帧维持某种层叠状态。

6.4 配饰

耳环、项链、眼镜、手表、发夹、戒指。所有可见的配饰都应被命名一次。配饰漂移很微妙 —— 第 60 帧手表消失、第 90 帧耳环不见 —— 但观众潜意识里最先察觉的恰恰是这些细节。

6.5 光线连续性

同一场景内的光线应当在 @tag 块或场景设置里明确。多场景之间保持光线连续更难 —— 但你可以锁定一个 方向(始终是相机右侧的轮廓光)和 色温(始终 5600K)。这能防止模型在场景推进时漂移到不同的情绪。

6.6 图案与质感保持

条纹、Logo、亮片、刺绣、织物纹路。这些都属于高漂移元素,因为模型倾向于 规则化 图案。在场景描述中要再次强调(”亮片捕捉到聚光灯”、”Logo 在同一位置保持清晰”)。

6.7 5 秒法则

如果场景描述里第 90 帧时主体没有变化,就再写一次。“同样的夹克,同样的颜色,没有图案漂移。” 这在 token 上只有很小的冗余成本,却能显著降低漂移。

关于锁定身份提示词写法的更深入讨论,请参阅 2026 年最佳 AI 视频提示词大师课。


7. 常见错误

7.1 主体描述过于模糊

“一个女人”或”一个穿夹克的人”等于没有给模型任何可锁定信息。每少写一个描述,模型就多出一个自由度,而每一帧它都会填上不同的答案。

7.2 互换 @tag 分配

如果 @character1 在场景 1 中引用、@character2 在场景 2 中引用,但两者的描述恰好互换,模型就会把身份互相对调。标签分配是身份契约的一部分。

7.3 风格线索互相冲突

在同一段提示词中同时写”真实感光照”和”动漫风格头发”不是简单的画面描述,而是反复触发的漂移源头。选择单一视觉风格并保持一致。对于风格迁移模式,应使用专门的 @style 参考图,而非在线索中混搭。

7.4 描述模型无法完成的运动

Ref2V 不会凭空赋予超越基础视频模型本身的运动能力。如果你要求一段 12 秒、180° 弧线的跟踪镜头,其实是在要求底层模型做它可能根本做不到的事。运动指令必须与底层模型文档承诺的能力对齐。

7.5 沿用 I2V 的提示词习惯

只写动作和场景(”镜头平移,她微笑”),而不指明哪些东西不能变,这是从 I2V 迁移过来最常见的错误。如果场景里没有写明身份,模型就无从锁定。

7.6 忽略音频参考

Seedance 2.5 的音频参考(最多 10 个)属于同一套 @tag 系统。如果你想在多场景作品中保持人声连续,只需要在场景中持续引用 @vo_main。音频漂移和视觉漂移一样真实存在。


8. 常见问题

Ref2V 和 I2V 在实操中的区别是什么?

在 I2V 中,源图就是第一帧,模型可以自由演化主体。在 Ref2V 中,参考图是贯穿每一帧的持久身份锚点,文本则描述围绕它发生的事。简单来说:Ref2V 让面部、产品和服装保持稳定;I2V 则允许它们变形。

Ref2V 能用多张参考图吗?

可以。Seedance 2.5 单次生成支持最多 30 张图片参考,每张都对应自己的 @tag。多主体合奏(双人角色、多产品)是 Ref2V 的标准模式之一。

视频和音频参考各能用多少?

视频参考最多 10 个,音频参考最多 10 个,加上 30 张图片参考 —— 单次生成总计 50 个多模态参考素材。

提示词最长能写多少?

Seedance 2.5 的提示词可以轻松容纳数百个 token,包括 @tag 定义和场景描述。5 秒法则(在第 90 帧再次描述保持不变的内容)完全在预算范围内。

Ref2V 能保证零漂移吗?

不能。相比 I2V,Ref2V 大幅降低 漂移,是当前最强的身份一致性方案,但 2026 年没有任何生产级模型能保证长片段零漂移。本指南中的防漂移清单能帮你最大程度接近零漂移。

可以把 Ref2V 和风格迁移结合起来吗?

可以 —— 这就是第三种参考模式。用 @style 参考图承载视觉风格,用 @subject 参考图承载身份,再用文本描述动作。避免在线索中把风格和身份混在一起写。

Ref2V 跟其他模型里的人物一致性工具是同一回事吗?

底层 目标 一致 —— 让主体在帧间保持稳定。实现机制 不同。Ref2V 是 arXiv:2508.02458 论文中正式化的训练范式;而 Kling、Runway、Sora 等模型中的人物一致性是通过不同的架构选择实现的。提示词的原则可以互通。

在哪里能找到更多提示词模板?

我们的 2K 视频模型提示词模板 和 Seedance 2.5 提示词指南 是最直接相关的延伸阅读。


9. 结语

Ref2V 是一次范式跃迁,把 AI 视频从尝鲜循环改造成真正的生产流水线。”一张参考图 + 目标文本”的格式在 arXiv:2508.02458 中正式化,并在 Seedance 2.5 的 50 参考素材系统中达到生产规模,让创作者能够把一张脸、一个产品或一种风格锁定到故事需要的任意多片段上。

本指南中的模板只是起点,让它们真正可用于生产的是防漂移技巧。把每一个 @tag 视为一份身份契约,把每一段场景描述同时视为动作契约 和 对”什么不能变”的再次声明。这种纪律正是把零散片段与成系列作品区分开的关键。

延伸阅读:Seedance 2.5 提示词指南、Kling 3.0 人物一致性详解、2K 视频模型提示词模板 以及 2026 年最佳 AI 视频提示词大师课。

参考资料


由 videosprompt.org 编辑团队审校 · 2026 年 10 月

分享文章

相关文章

推荐阅读

开始你的下一步

探索更多可能,发现适合你的解决方案。