参考图生成视频模板:Seedance、Veo、Kling、Wan 及更多模型的跨模型实战手册
TL;DR
- 参考图生成视频(Ref2V) 使用一张或多张静态图像作为生成视频片段的视觉锚点,而不是仅依赖文本。
- 三种核心 Ref2V 技术分别是单主体参考、多主体参考和仅风格参考,每种都有不同的脚手架和失败模式。
- 本文提供 12 份即用模板,适配 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0。
- 一套通用 Ref2V 骨架——锚点、主体令牌、场景、运动/镜头、音频——可在全部四个模型 API 间复用。
- MiniMax H3 仅被提及,未经实测验证。 我们尚未将 H3 投入这些模板的实测;本手册基于四个有据可查的 API 进行验证,并附带 H3 的验证清单。
- 仅风格参考是最容易验证、成本最低的方式,也是任何新模型(包括 H3)的最佳起点。
- 文末的验证流程允许你对任何模型(包括 H3)按同一套五组件骨架进行打分。
1. 2026 年”参考图生成视频”的含义
参考图生成视频(Ref2V)以一张或多张输入图像为条件约束视频模型,使输出继承特定的视觉属性——面部、产品、光照风格、品牌色调——而这些是单纯文本无法可靠描述的。Ref2V 论文(arXiv:2508.02458)正式确立了主体令牌交接(subject-token handoff)模式,所有主流商用 API 现在都以不同名称暴露该模式。
Ref2V 内部包含三种核心技术:
| 技术 | 主要用途 | 输入 | 失败模式 |
|---|---|---|---|
| 单主体参考 | 在整个片段中锁定单一身份(面部、产品、角色) | 1 张参考图 + 场景描述 | 当场景提示压过参考时,出现身份漂移 |
| 多主体参考 | 将两个或更多角色/对象合成到同一场景中 | 2–N 张参考图 + 角色标签 | 当令牌未被消歧时,出现主体串扰和姿态混淆 |
| 仅风格参考 | 继承光照、色彩或艺术风格处理 | 1 张参考图 + 仅风格提示 | 风格压过主体;主体被重新风格化而非被重新渲染 |
区分这三种技术至关重要,因为每种技术有不同的提示骨架、成本特征和失败风险。一份实用的 Ref2V 手册需要同时处理三者,而非将它们混为一谈。
2. 我们验证了什么,未能验证什么
已针对有据可查的 API 进行验证。 这 12 份模板基于四个模型的有公开文档的接口进行了适配:
- Seedance 2.5——
@tag参考语法和 50 素材工作流(Seedance 50 参考图指南)。 - Veo 3.1——
image_prompt字段(deepmind.google/models/veo)。 - Kling 3.0——用于多参考图绑定的
elements数组。 - Wan 3.0——
start_image和end_image字段。
有关 prompt 工程背景,请参阅 mstudio.ai 影视制作提示指南。
MiniMax H3 仅被提及,未经实测验证。 本文的关键词涉及 H3,因此我们在标题中包含它以提升可发现性。我们尚未将 H3 投入这些模板的实测,也不清楚 H3 使用的是 @tag 风格语法、image_prompt、elements[],还是其他机制。请将这些模板作为假设施加于 H3,并使用第 8 节对实际输出进行打分。
我们不会凭空编造 H3 专属的字段名,也不会声称未跑过的基准测试数据。
3. Ref2V 提示骨架
每个 Ref2V 提示都可拆解为五个组件。我们在全部 12 份模板中都使用这套骨架,因为它在不同 API 之间移植时比各模型原生语法更经得起考验。
| # | 组件 | 用途 | 通用示例 |
|---|---|---|---|
| 1 | 来源参考锚点 | 告知模型应查阅哪些图像 | reference_image: ./subject.png |
| 2 | 主体令牌描述 | 命名锚点所指对象,以便模型重新绑定 | subject_token: "the woman in the red coat" |
| 3 | 场景内容 | 描述主体被放入的新场景 | scene: "walking through a Tokyo alley at night, neon reflections on wet pavement" |
| 4 | 运动/镜头方向 | 指定场景如何运动 | camera: "slow dolly-in, eye-level, 24fps cinematic" |
| 5 | 音频线索 | 设定音效设计(或静音) | audio: "rain ambience, distant city hum, footsteps" |
这套骨架之所以可移植,是因为每家商用 Ref2V API 都暴露了这五个槽位,即便名称各异。Seedance 将锚点包装在 @tag 中,Veo 使用 image_prompt,Kling 使用带角色键的 elements[],Wan 使用 start_image/end_image。五组件始终如一。
请将主体令牌描述保持简短——2 到 6 个单词。过长的自然语言主体描述会重新引入 Ref2V 本应消除的歧义。
4. 各模型模板适配
下面展示同一 Ref2V 意图如何在有据可查的 API 之间转换,以一个单主体示例为例:一位身穿红色外套的女子在夜晚穿过东京小巷。
| 组件 | Seedance 2.5 | Veo 3.1 | Kling 3.0 | Wan 3.0 |
|---|---|---|---|---|
| 来源锚点 | @subject(附带图像) |
image_prompt: <image> |
elements[0].image |
start_image: <image> |
| 主体令牌 | the woman in the red coat |
在 prompt 字段中描述 |
elements[0].description |
在 prompt 字段中描述 |
| 场景内容 | 内联于 prompt 中 | prompt 字段 |
prompt 字段 |
prompt 字段 |
| 运动/镜头 | 在 // camera: 之后内联 |
内联于 prompt 中 |
内联于 prompt 中 |
end_image + 镜头内联于 prompt |
| 音频线索 | 内联 // audio: |
audio_prompt 字段 |
audio 字段 |
未暴露;占位处理 |
三点观察:
- Seedance 和 Veo 表达力最强。Seedance 的
@tag系统在多参考场景下最为便捷;Veo 的image_prompt是最简洁的单主体接口。 - Kling 的
elements数组 是最显式的角色绑定机制——在多主体参考(主体串扰是主要风险)场景下最为强劲。 - Wan 的首尾帧方案 限制最多——必须指定首帧和尾帧,这对运动驱动的参考是优势,但对自由形式的 Ref2V 则是局限。
本表是整份手册的脊梁。当某份模板未明确指定模型时,沿用通用骨架,并参照本表做适配。
5. 12 份提示模板
每份模板都是一个围栏 text 代码块。针对各模型的适配方案见第 6 节。
5.1 单主体参考(3 份模板)
模板 1——行走人像(单主体,城市)
# Ref2V — Single-subject walking portrait
# Identity-locked B-roll
[anchor]
reference_image: ./woman_red_coat.png
[subject]
subject_token: "the woman in the red coat"
preserve: face, hair length, coat color, coat cut
[scene]
location: Tokyo, Shinjuku alley, night
environment: wet pavement, neon (blue and magenta), light rain, steam from vents
time: 23:40
[motion]
camera: slow dolly-in, eye-level, 35mm feel
subject_motion: walks toward camera, slight head turn at the end
duration: 6s
fps: 24
[audio]
ambience: rain on pavement, distant traffic
foley: footsteps, coat rustle
模板 2——产品展示(单主体,工作室)
# Ref2V — Single-subject product showcase
# E-commerce, brand lock
[anchor]
reference_image: ./ceramic_vase.png
[subject]
subject_token: "the matte-black ceramic vase"
preserve: silhouette, glaze finish, base footprint
[scene]
location: studio, seamless off-white backdrop
environment: soft key light from camera-left, subtle rim from right
[motion]
camera: slow 180-degree turntable, slight push-in at midpoint
subject_motion: static, camera orbits
duration: 5s
fps: 30
[audio]
ambience: silent room tone
模板 3——角色对白节拍(单主体,表演)
# Ref2V — Single-subject dialogue beat
# Narrative, emotional close-up
[anchor]
reference_image: ./character_portrait.png
[subject]
subject_token: "the man with the grey beard"
preserve: face, beard shape, eye color
[scene]
location: dimly lit diner booth
environment: warm tungsten key, cool window fill
time: 02:15
[motion]
camera: locked medium close-up, very subtle handheld drift
subject_motion: slight head tilt down, then up; eyes meet lens at the end
duration: 4s
fps: 24
[audio]
ambience: diner kitchen clatter, soft jazz underscore
foley: shallow breathing
5.2 多主体参考(3 份模板)
模板 4——双人对话
# Ref2V — Two-character conversation
# Dialogue scene with identity preservation
[anchors]
reference_image_a: ./character_A.png
reference_image_b: ./character_B.png
[subjects]
subject_token_a: "the woman with the short black hair"
subject_token_b: "the man in the grey cardigan"
preserve: face, hair, primary garment color per subject
[scene]
location: park bench, autumn afternoon
environment: golden-hour side key, leaf fall, clear and cool
time: 17:45
[motion]
camera: slow push-in from wide to medium two-shot
subject_motion: A gestures with right hand while speaking; B nods, looks down, then back up
duration: 7s
fps: 24
[audio]
ambience: distant children, wind in leaves
foley: rustling jacket, bench creak
模板 5——主体与道具交接
# Ref2V — Subject + prop handoff
# Brand + actor, unboxing narrative
[anchors]
reference_image_subject: ./presenter.png
reference_image_prop: ./product_box.png
[subjects]
subject_token_subject: "the presenter in the navy blazer"
subject_token_prop: "the teal product box with the silver logo"
preserve: subject face and blazer cut; prop dimensions, logo placement, color saturation
[scene]
location: clean studio set
environment: high-key soft lighting, gradient grey backdrop
[motion]
camera: static medium shot, focus pull subject → prop at handoff
subject_motion: subject lifts box from table with both hands, holds chest-high, looks down; prop held still, slight reflection shift
duration: 5s
fps: 30
[audio]
ambience: silent
foley: cardboard slide on table
music: corporate underscore, fades in
模板 6——群戏阵容(三个主体)
# Ref2V — Ensemble cast
# Group shot, foreground/background separation
[anchors]
reference_image_a: ./cast_A.png
reference_image_b: ./cast_B.png
reference_image_c: ./cast_C.png
[subjects]
subject_token_a: "the older man with the round glasses"
subject_token_b: "the young woman with the red scarf"
subject_token_c: "the child in the denim jacket"
preserve: face and signature accessory per subject
[scene]
location: subway platform, midday
environment: fluorescent overhead, daylight wall behind
time: 12:10
[motion]
camera: static wide shot
subject_motion: a stands left, hands in pockets, slight sway; b stands center, scrolls phone; c stands right, looks up tunnel, points
duration: 6s
fps: 24
[audio]
ambience: train rumble, station announcement (indistinct)
foley: footsteps, paper rustle
5.3 仅风格参考(3 份模板)
模板 7——电影级色彩调色
# Ref2V — Style-only: cinematic color grade
# Lock a teal-and-orange look
[anchor]
reference_image: ./reference_grade.png
[subject]
preserve: none (style-only)
[style]
inherit:
palette: teal shadows, orange highlights, crushed blacks
contrast: high
grain: medium 35mm
halation: subtle warm bloom on highlights
do_not_inherit:
composition: ignore layout of reference
subjects: do not copy figures from reference
[scene]
scene_subject: a courier on a motorcycle riding through city traffic at dusk
[motion]
camera: tracking shot, slightly low angle
duration: 5s
fps: 24
[audio]
ambience: traffic hum
模板 8——绘画风风格迁移
# Ref2V — Style-only: painterly style
# Art-direction lock
[anchor]
reference_image: ./painting_reference.png
[subject]
preserve: scene composition, but render as if painted
[style]
inherit:
medium: oil on canvas
brushwork: visible, loose in background, tighter on focal subject
palette: muted earth tones with one accent (the painting's accent)
texture: canvas grain visible throughout
do_not_inherit:
subjects: do not copy figures; re-render new subjects in the same style
[scene]
scene_subject: a baker arranging bread in a window display, morning light
[motion]
camera: locked medium shot
duration: 4s
fps: 24
[audio]
ambience: city street outside
模板 9——品牌视觉识别
# Ref2V — Style-only: brand visual identity
# Campaign footage matching brand deck
[anchor]
reference_image: ./brand_moodboard.png
[subject]
preserve: brand palette, typography cue, motion cadence
[style]
inherit:
palette: brand primary (deep indigo) + accent (warm gold), off-white negative
composition: generous negative space, subject lower-third
motion: slow, deliberate, 1.2x slower than real time
text_treatment: leave room for headline overlay (lower-third safe area)
do_not_inherit:
subjects: do not copy figures from moodboard
props: do not copy specific products
[scene]
scene_subject: a small business owner unlocking their shop door at sunrise
[motion]
camera: static wide, then slow push-in during final second
duration: 6s
fps: 30
[audio]
ambience: morning birds
music: brand sonic logo sting at the end (placeholder)
5.4 运动驱动参考(3 份模板)
模板 10——首尾帧运动
# Ref2V — Motion-driven: start/end frame
# Wan-style first/last frame anchoring
[anchors]
start_image: ./frame_start.png # standing still
end_image: ./frame_end.png # mid-stride, arm raised
[subject]
subject_token: "the woman in the olive jacket"
preserve: face, jacket, bag strap
[scene]
location: city crosswalk, daytime, dry
environment: bright overcast, soft shadows
time: 14:30
[motion]
camera: locked eye-level medium shot
subject_motion: interpolates from standing still to mid-stride with raised arm
duration: 3s
fps: 30
[audio]
ambience: crosswalk signal, distant traffic
foley: footsteps
模板 11——纯镜头运动
# Ref2V — Motion-driven: camera-only
# Parallax reveal, locked subject
[anchor]
reference_image: ./interior_room.png
[subject]
subject_token: "the room interior (no specific human subject)"
preserve: composition, furniture placement, lighting
[scene]
location: loft apartment, late afternoon
environment: window light from camera-right
time: 18:10
[motion]
camera: slow dolly from doorway toward window, slight arc
subject_motion: none (subject is the room itself)
duration: 7s
fps: 24
[audio]
ambience: city hum outside window
music: ambient pad, low volume
模板 12——带主体运动的动作弧线
# Ref2V — Motion-driven: action arc
# Dynamic beat with identity lock
[anchor]
reference_image: ./dancer.png
[subject]
subject_token: "the dancer in the white shirt"
preserve: face, shirt, hairstyle
[scene]
location: empty studio, single spotlight from above
environment: black backdrop, dust in the spotlight cone
[motion]
camera: handheld, orbits the subject in a half-circle during the move
subject_motion: leaps from a crouch, spins once mid-air, lands in a wide stance
duration: 4s
fps: 60
[audio]
ambience: silent studio
foley: shoe squeak, landing thud
music: none (post)
6. 12 份模板在各模型间的适配
下表展示模板 1(行走人像)转换为各模型 API 接口的版本。共享场景文本(下方以 <scene> 引用):”the woman in the red coat walks toward camera through a Tokyo alley at night. Wet pavement reflects blue and magenta neon. Light rain. Steam from vents.” 共享镜头:”slow dolly-in, eye-level, 35mm cinematic.”
Seedance 2.5
@subject: ./woman_red_coat.png
<scene>
camera: slow dolly-in, eye-level, 35mm cinematic feel.
// audio: rain on pavement, distant traffic, footsteps, coat fabric
@subject 将图像绑定到最近的名称短语;// 注释会被忽略;多主体请加 @subject2、@subject3。
Veo 3.1
{
"image_prompt": "./woman_red_coat.png",
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic",
"audio_prompt": "rain on pavement, distant traffic, footsteps, coat fabric"
}
image_prompt 是锚点;其余内容放入文本字段。场景描述请控制在约 120 词以内。
Kling 3.0
{
"elements": [
{"image": "./woman_red_coat.png", "description": "the woman in the red coat", "role": "primary_subject"}
],
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic",
"audio": "rain on pavement, distant traffic, footsteps, coat fabric"
}
elements 是最显式的角色绑定接口。如为仅风格模式,请设置 role: "style_reference" 并将 description 留空。
Wan 3.0
{
"start_image": "./woman_red_coat.png",
"end_image": "./woman_red_coat_frame2.png",
"prompt": "<scene>",
"camera": "slow dolly-in, eye-level, 35mm cinematic"
}
Wan 在运动驱动参考下要求同时提供 start_image 和 end_image。音频未原生暴露,请在后期添加。
全部 12 份模板的适配矩阵
| 模板 | Seedance 2.5 | Veo 3.1 | Kling 3.0 | Wan 3.0 | 说明 |
|---|---|---|---|---|---|
| 1. 行走人像 | @subject |
image_prompt |
elements[0] |
start_image+end_image |
Wan 需手绘尾帧 |
| 2. 产品展示 | @subject+转台 |
image_prompt |
elements[0] |
start_image |
四者皆可处理 |
| 3. 对白节拍 | @subject |
image_prompt |
elements[0] |
适配度弱 | Wan 难以捕捉微妙表情 |
| 4. 双人对话 | @subject+@subject2 |
两个 image_prompt |
elements[0..1] |
不支持 | Kling 胜出 |
| 5. 主体+道具 | @subject+@prop |
两个 image_prompt |
elements[0..1] |
适配度弱 | Seedance/Kling 强劲 |
| 6. 群戏阵容 | @subject ×3 |
不实用 | elements[0..2] |
不支持 | 仅 Kling 一次性可成 |
| 7. 电影级调色 | @style |
image_prompt+线索 |
elements[0] style |
start_image+线索 |
四者皆可处理风格 |
| 8. 绘画风风格 | @style |
image_prompt+媒介 |
elements[0] style |
start_image |
四者皆可处理 |
| 9. 品牌识别 | @style+品牌线索 |
image_prompt+线索 |
elements[0] style |
start_image+线索 |
四者皆可处理 |
| 10. 首尾帧 | 非原生 | 非原生 | 非原生 | 原生 | Wan 胜出 |
| 11. 纯镜头运动 | @scene+camera |
image_prompt+camera |
elements[0]+camera |
适配度弱 | Veo/Seedance 强劲 |
| 12. 动作弧线 | @subject+运动 |
image_prompt+运动 |
elements[0]+运动 |
两帧法 | 四者皆可,需注意细节 |
7. 姊妹文章与延伸阅读
- Ref2V 配套文章——Seedance Ref2V 详解。
- Seedance 2.5 提示指南——Seedance 2.5 模型指南。
- Seedance 多参考图提示——50 素材配套文章。
- 图生视频角色锁定提示——相邻的 I2V 方向。
- 多模态参考视频提示——多模态上下文。
- 2026 年最佳 AI 视频提示大师课——跨模型概览。
外部资料:
8. 验证 H3 兼容性
一种可复现的方法,按五组件骨架对 H3 进行打分:
确认 H3 接受哪些输入。 检查 H3 API 中是否存在参考图像字段(
image_prompt、reference_image、init_image、@tag,或elements[])、主体令牌机制、仅风格支持,以及首尾帧支持。若无法同时确认四项,则将 H3 视为部分已定义。从模板 7(电影级调色)开始。 仅风格模式验证最易。无法通过仅风格测试的模型将无法通过其他任何测试。
运行模板 1(行走人像)。 这是规范的单主体 Ref2V 试验。
运行模板 4(双人对话)。 这是首个会暴露主体串扰的测试。若 H3 出现面部互换或特征合并,即为明确的失败信号。
按五个维度为每次输出打分(0 = 未遵循,1 = 部分遵循,2 = 完全遵循):
| 维度 | 检查内容 | |———–|—————| | 锚点保真度 | 主体是否与参考图相似? | | 风格保真度 | (仅风格模式)风格是否匹配? | | 场景遵循度 | 场景是否与文字描述一致? | | 运动遵循度 | 运动是否符合镜头/主体指令? | | 音频遵循度 | 音频是否符合(如支持)? |
10 分中得 8 分及以上,表示 H3 与通用骨架兼容;低于 8 分则意味着 H3 需要自己的适配方案。
- 记录失败。 某个组件失败(例如音频)属于功能缺失,而非对 H3 整体的判决。记下失败点,并使用不依赖该组件的模板。
FAQ
什么是参考图生成视频提示模板?
参考图生成视频(Ref2V)提示模板将参考图像与结构化文本搭配,使视频模型生成的输出继承参考图的视觉属性——身份、风格、品牌或运动。规范结构包含五个组件:锚点、主体令牌、场景、运动、音频。请参阅 arXiv:2508.02458。
Ref2V 与 I2V 有何不同?
I2V 动画化单一来源图像并预测合理的运动。Ref2V 将输入视为*参考*——对身份、风格或构图的约束——并与定义新场景的文本配对。当你需要身份锁定或品牌保真时使用 Ref2V;当你只想让静态图像动起来时使用 I2V。
哪个模型最适合多主体参考?
Kling 3.0 通过其带逐主体角色键的 elements 数组实现。Seedance 2.5 通过 @subject 标签支持,是最便捷的内联方案。Veo 3.1 需要使用多个 image_prompt 字段并合成。Wan 3.0 在单次调用中不支持多主体。
我可以将这些模板用于 H3 或 S2V/V2V 吗?
本模板基于有据可查的 API 设计,尚未针对 H3 进行验证;字段名称可能有所不同。五组件骨架应当可以平移,但请按第 8 节对 H3 打分。本模板专门针对 Ref2V——S2V 与 V2V 相关,但语法不可直接复用。
这篇文章适用于 H3 吗?
可以,作为*起点*。我们未对 H3 进行测试。请将模板视为假设,按第 8 节运行测试,并按 H3 的实际 API 接口调整语法。获得独立 H3 数据后我们将更新本文。
Ref2V 提示应该多长?
场景描述请控制在约 120 词以内,主体令牌控制在 2–6 词。过长的描述会重新引入 Ref2V 本应消除的歧义。
结论
2026 年的参考图生成视频提示是一套五组件骨架,可跨 Seedance、Veo、Kling、Wan 复用——一旦经过验证,也可能适用于 H3 及其他即将推出的模型。请将 12 份模板视为起点,按第 8 节为你的输出打分,并根据你所对接的 API 接口调整语法。
Seedance 深入内容,请参阅 Seedance 2.5 提示指南和多参考图 50 素材配套文章。多模态框架相关内容,请参阅多模态参考视频提示。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
本文局限性
本 12 份模板已基于 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0 的有据可查 API 进行验证。H3 因关键词可发现性需要而出现在标题与 slug 中,我们尚未独立地将 H3 投入这些模板的实测,也缺乏关于其参考图像 API 接口的第一手数据。将其应用于 H3 的读者应将模板视为假设,并按第 8 节操作。H3 上的结果可能与有据可查的模型有所不同。获得独立 H3 数据后我们将修订本文。
分享文章