VideosPrompt VideosPrompt

参考图生成视频模板:Seedance、Veo、Kling、Wan 及更多模型的跨模型实战手册

作者: VideosPrompt 日期: 2026-10-07 14:12:04
参考图生成视频模板:Seedance、Veo、Kling、Wan 及更多模型的跨模型实战手册

TL;DR

  • 参考图生成视频(Ref2V) 使用一张或多张静态图像作为生成视频片段的视觉锚点,而不是仅依赖文本。
  • 三种核心 Ref2V 技术分别是单主体参考、多主体参考和仅风格参考,每种都有不同的脚手架和失败模式。
  • 本文提供 12 份即用模板,适配 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0。
  • 一套通用 Ref2V 骨架——锚点、主体令牌、场景、运动/镜头、音频——可在全部四个模型 API 间复用。
  • MiniMax H3 仅被提及,未经实测验证。 我们尚未将 H3 投入这些模板的实测;本手册基于四个有据可查的 API 进行验证,并附带 H3 的验证清单。
  • 仅风格参考是最容易验证、成本最低的方式,也是任何新模型(包括 H3)的最佳起点。
  • 文末的验证流程允许你对任何模型(包括 H3)按同一套五组件骨架进行打分。

1. 2026 年”参考图生成视频”的含义

参考图生成视频(Ref2V)以一张或多张输入图像为条件约束视频模型,使输出继承特定的视觉属性——面部、产品、光照风格、品牌色调——而这些是单纯文本无法可靠描述的。Ref2V 论文(arXiv:2508.02458)正式确立了主体令牌交接(subject-token handoff)模式,所有主流商用 API 现在都以不同名称暴露该模式。

Ref2V 内部包含三种核心技术:

技术 主要用途 输入 失败模式
单主体参考 在整个片段中锁定单一身份(面部、产品、角色) 1 张参考图 + 场景描述 当场景提示压过参考时,出现身份漂移
多主体参考 将两个或更多角色/对象合成到同一场景中 2–N 张参考图 + 角色标签 当令牌未被消歧时,出现主体串扰和姿态混淆
仅风格参考 继承光照、色彩或艺术风格处理 1 张参考图 + 仅风格提示 风格压过主体;主体被重新风格化而非被重新渲染

区分这三种技术至关重要,因为每种技术有不同的提示骨架、成本特征和失败风险。一份实用的 Ref2V 手册需要同时处理三者,而非将它们混为一谈。


2. 我们验证了什么,未能验证什么

已针对有据可查的 API 进行验证。 这 12 份模板基于四个模型的有公开文档的接口进行了适配:

  • Seedance 2.5——@tag 参考语法和 50 素材工作流(Seedance 50 参考图指南)。
  • Veo 3.1——image_prompt 字段(deepmind.google/models/veo)。
  • Kling 3.0——用于多参考图绑定的 elements 数组。
  • Wan 3.0——start_image 和 end_image 字段。

有关 prompt 工程背景,请参阅 mstudio.ai 影视制作提示指南。

MiniMax H3 仅被提及,未经实测验证。 本文的关键词涉及 H3,因此我们在标题中包含它以提升可发现性。我们尚未将 H3 投入这些模板的实测,也不清楚 H3 使用的是 @tag 风格语法、image_prompt、elements[],还是其他机制。请将这些模板作为假设施加于 H3,并使用第 8 节对实际输出进行打分。

我们不会凭空编造 H3 专属的字段名,也不会声称未跑过的基准测试数据。


3. Ref2V 提示骨架

每个 Ref2V 提示都可拆解为五个组件。我们在全部 12 份模板中都使用这套骨架,因为它在不同 API 之间移植时比各模型原生语法更经得起考验。

# 组件 用途 通用示例
1 来源参考锚点 告知模型应查阅哪些图像 reference_image: ./subject.png
2 主体令牌描述 命名锚点所指对象,以便模型重新绑定 subject_token: "the woman in the red coat"
3 场景内容 描述主体被放入的新场景 scene: "walking through a Tokyo alley at night, neon reflections on wet pavement"
4 运动/镜头方向 指定场景如何运动 camera: "slow dolly-in, eye-level, 24fps cinematic"
5 音频线索 设定音效设计(或静音) audio: "rain ambience, distant city hum, footsteps"

这套骨架之所以可移植,是因为每家商用 Ref2V API 都暴露了这五个槽位,即便名称各异。Seedance 将锚点包装在 @tag 中,Veo 使用 image_prompt,Kling 使用带角色键的 elements[],Wan 使用 start_image/end_image。五组件始终如一。

请将主体令牌描述保持简短——2 到 6 个单词。过长的自然语言主体描述会重新引入 Ref2V 本应消除的歧义。


4. 各模型模板适配

下面展示同一 Ref2V 意图如何在有据可查的 API 之间转换,以一个单主体示例为例:一位身穿红色外套的女子在夜晚穿过东京小巷。

组件 Seedance 2.5 Veo 3.1 Kling 3.0 Wan 3.0
来源锚点 @subject(附带图像) image_prompt: <image> elements[0].image start_image: <image>
主体令牌 the woman in the red coat 在 prompt 字段中描述 elements[0].description 在 prompt 字段中描述
场景内容 内联于 prompt 中 prompt 字段 prompt 字段 prompt 字段
运动/镜头 在 // camera: 之后内联 内联于 prompt 中 内联于 prompt 中 end_image + 镜头内联于 prompt
音频线索 内联 // audio: audio_prompt 字段 audio 字段 未暴露;占位处理

三点观察:

  1. Seedance 和 Veo 表达力最强。Seedance 的 @tag 系统在多参考场景下最为便捷;Veo 的 image_prompt 是最简洁的单主体接口。
  2. Kling 的 elements 数组 是最显式的角色绑定机制——在多主体参考(主体串扰是主要风险)场景下最为强劲。
  3. Wan 的首尾帧方案 限制最多——必须指定首帧和尾帧,这对运动驱动的参考是优势,但对自由形式的 Ref2V 则是局限。

本表是整份手册的脊梁。当某份模板未明确指定模型时,沿用通用骨架,并参照本表做适配。


5. 12 份提示模板

每份模板都是一个围栏 text 代码块。针对各模型的适配方案见第 6 节。

5.1 单主体参考(3 份模板)

模板 1——行走人像(单主体,城市)

# Ref2V — Single-subject walking portrait
# Identity-locked B-roll

[anchor]
reference_image: ./woman_red_coat.png

[subject]
subject_token: "the woman in the red coat"
preserve: face, hair length, coat color, coat cut

[scene]
location: Tokyo, Shinjuku alley, night
environment: wet pavement, neon (blue and magenta), light rain, steam from vents
time: 23:40

[motion]
camera: slow dolly-in, eye-level, 35mm feel
subject_motion: walks toward camera, slight head turn at the end
duration: 6s
fps: 24

[audio]
ambience: rain on pavement, distant traffic
foley: footsteps, coat rustle

模板 2——产品展示(单主体,工作室)

# Ref2V — Single-subject product showcase
# E-commerce, brand lock

[anchor]
reference_image: ./ceramic_vase.png

[subject]
subject_token: "the matte-black ceramic vase"
preserve: silhouette, glaze finish, base footprint

[scene]
location: studio, seamless off-white backdrop
environment: soft key light from camera-left, subtle rim from right

[motion]
camera: slow 180-degree turntable, slight push-in at midpoint
subject_motion: static, camera orbits
duration: 5s
fps: 30

[audio]
ambience: silent room tone

模板 3——角色对白节拍(单主体,表演)

# Ref2V — Single-subject dialogue beat
# Narrative, emotional close-up

[anchor]
reference_image: ./character_portrait.png

[subject]
subject_token: "the man with the grey beard"
preserve: face, beard shape, eye color

[scene]
location: dimly lit diner booth
environment: warm tungsten key, cool window fill
time: 02:15

[motion]
camera: locked medium close-up, very subtle handheld drift
subject_motion: slight head tilt down, then up; eyes meet lens at the end
duration: 4s
fps: 24

[audio]
ambience: diner kitchen clatter, soft jazz underscore
foley: shallow breathing

5.2 多主体参考(3 份模板)

模板 4——双人对话

# Ref2V — Two-character conversation
# Dialogue scene with identity preservation

[anchors]
reference_image_a: ./character_A.png
reference_image_b: ./character_B.png

[subjects]
subject_token_a: "the woman with the short black hair"
subject_token_b: "the man in the grey cardigan"
preserve: face, hair, primary garment color per subject

[scene]
location: park bench, autumn afternoon
environment: golden-hour side key, leaf fall, clear and cool
time: 17:45

[motion]
camera: slow push-in from wide to medium two-shot
subject_motion: A gestures with right hand while speaking; B nods, looks down, then back up
duration: 7s
fps: 24

[audio]
ambience: distant children, wind in leaves
foley: rustling jacket, bench creak

模板 5——主体与道具交接

# Ref2V — Subject + prop handoff
# Brand + actor, unboxing narrative

[anchors]
reference_image_subject: ./presenter.png
reference_image_prop: ./product_box.png

[subjects]
subject_token_subject: "the presenter in the navy blazer"
subject_token_prop: "the teal product box with the silver logo"
preserve: subject face and blazer cut; prop dimensions, logo placement, color saturation

[scene]
location: clean studio set
environment: high-key soft lighting, gradient grey backdrop

[motion]
camera: static medium shot, focus pull subject → prop at handoff
subject_motion: subject lifts box from table with both hands, holds chest-high, looks down; prop held still, slight reflection shift
duration: 5s
fps: 30

[audio]
ambience: silent
foley: cardboard slide on table
music: corporate underscore, fades in

模板 6——群戏阵容(三个主体)

# Ref2V — Ensemble cast
# Group shot, foreground/background separation

[anchors]
reference_image_a: ./cast_A.png
reference_image_b: ./cast_B.png
reference_image_c: ./cast_C.png

[subjects]
subject_token_a: "the older man with the round glasses"
subject_token_b: "the young woman with the red scarf"
subject_token_c: "the child in the denim jacket"
preserve: face and signature accessory per subject

[scene]
location: subway platform, midday
environment: fluorescent overhead, daylight wall behind
time: 12:10

[motion]
camera: static wide shot
subject_motion: a stands left, hands in pockets, slight sway; b stands center, scrolls phone; c stands right, looks up tunnel, points
duration: 6s
fps: 24

[audio]
ambience: train rumble, station announcement (indistinct)
foley: footsteps, paper rustle

5.3 仅风格参考(3 份模板)

模板 7——电影级色彩调色

# Ref2V — Style-only: cinematic color grade
# Lock a teal-and-orange look

[anchor]
reference_image: ./reference_grade.png

[subject]
preserve: none (style-only)

[style]
inherit:
  palette: teal shadows, orange highlights, crushed blacks
  contrast: high
  grain: medium 35mm
  halation: subtle warm bloom on highlights
do_not_inherit:
  composition: ignore layout of reference
  subjects: do not copy figures from reference

[scene]
scene_subject: a courier on a motorcycle riding through city traffic at dusk

[motion]
camera: tracking shot, slightly low angle
duration: 5s
fps: 24

[audio]
ambience: traffic hum

模板 8——绘画风风格迁移

# Ref2V — Style-only: painterly style
# Art-direction lock

[anchor]
reference_image: ./painting_reference.png

[subject]
preserve: scene composition, but render as if painted

[style]
inherit:
  medium: oil on canvas
  brushwork: visible, loose in background, tighter on focal subject
  palette: muted earth tones with one accent (the painting's accent)
  texture: canvas grain visible throughout
do_not_inherit:
  subjects: do not copy figures; re-render new subjects in the same style

[scene]
scene_subject: a baker arranging bread in a window display, morning light

[motion]
camera: locked medium shot
duration: 4s
fps: 24

[audio]
ambience: city street outside

模板 9——品牌视觉识别

# Ref2V — Style-only: brand visual identity
# Campaign footage matching brand deck

[anchor]
reference_image: ./brand_moodboard.png

[subject]
preserve: brand palette, typography cue, motion cadence

[style]
inherit:
  palette: brand primary (deep indigo) + accent (warm gold), off-white negative
  composition: generous negative space, subject lower-third
  motion: slow, deliberate, 1.2x slower than real time
  text_treatment: leave room for headline overlay (lower-third safe area)
do_not_inherit:
  subjects: do not copy figures from moodboard
  props: do not copy specific products

[scene]
scene_subject: a small business owner unlocking their shop door at sunrise

[motion]
camera: static wide, then slow push-in during final second
duration: 6s
fps: 30

[audio]
ambience: morning birds
music: brand sonic logo sting at the end (placeholder)

5.4 运动驱动参考(3 份模板)

模板 10——首尾帧运动

# Ref2V — Motion-driven: start/end frame
# Wan-style first/last frame anchoring

[anchors]
start_image: ./frame_start.png   # standing still
end_image: ./frame_end.png       # mid-stride, arm raised

[subject]
subject_token: "the woman in the olive jacket"
preserve: face, jacket, bag strap

[scene]
location: city crosswalk, daytime, dry
environment: bright overcast, soft shadows
time: 14:30

[motion]
camera: locked eye-level medium shot
subject_motion: interpolates from standing still to mid-stride with raised arm
duration: 3s
fps: 30

[audio]
ambience: crosswalk signal, distant traffic
foley: footsteps

模板 11——纯镜头运动

# Ref2V — Motion-driven: camera-only
# Parallax reveal, locked subject

[anchor]
reference_image: ./interior_room.png

[subject]
subject_token: "the room interior (no specific human subject)"
preserve: composition, furniture placement, lighting

[scene]
location: loft apartment, late afternoon
environment: window light from camera-right
time: 18:10

[motion]
camera: slow dolly from doorway toward window, slight arc
subject_motion: none (subject is the room itself)
duration: 7s
fps: 24

[audio]
ambience: city hum outside window
music: ambient pad, low volume

模板 12——带主体运动的动作弧线

# Ref2V — Motion-driven: action arc
# Dynamic beat with identity lock

[anchor]
reference_image: ./dancer.png

[subject]
subject_token: "the dancer in the white shirt"
preserve: face, shirt, hairstyle

[scene]
location: empty studio, single spotlight from above
environment: black backdrop, dust in the spotlight cone

[motion]
camera: handheld, orbits the subject in a half-circle during the move
subject_motion: leaps from a crouch, spins once mid-air, lands in a wide stance
duration: 4s
fps: 60

[audio]
ambience: silent studio
foley: shoe squeak, landing thud
music: none (post)

6. 12 份模板在各模型间的适配

下表展示模板 1(行走人像)转换为各模型 API 接口的版本。共享场景文本(下方以 <scene> 引用):”the woman in the red coat walks toward camera through a Tokyo alley at night. Wet pavement reflects blue and magenta neon. Light rain. Steam from vents.” 共享镜头:”slow dolly-in, eye-level, 35mm cinematic.”

Seedance 2.5

@subject: ./woman_red_coat.png
<scene>
camera: slow dolly-in, eye-level, 35mm cinematic feel.
// audio: rain on pavement, distant traffic, footsteps, coat fabric

@subject 将图像绑定到最近的名称短语;// 注释会被忽略;多主体请加 @subject2、@subject3。

Veo 3.1

{
  "image_prompt": "./woman_red_coat.png",
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic",
  "audio_prompt": "rain on pavement, distant traffic, footsteps, coat fabric"
}

image_prompt 是锚点;其余内容放入文本字段。场景描述请控制在约 120 词以内。

Kling 3.0

{
  "elements": [
    {"image": "./woman_red_coat.png", "description": "the woman in the red coat", "role": "primary_subject"}
  ],
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic",
  "audio": "rain on pavement, distant traffic, footsteps, coat fabric"
}

elements 是最显式的角色绑定接口。如为仅风格模式,请设置 role: "style_reference" 并将 description 留空。

Wan 3.0

{
  "start_image": "./woman_red_coat.png",
  "end_image": "./woman_red_coat_frame2.png",
  "prompt": "<scene>",
  "camera": "slow dolly-in, eye-level, 35mm cinematic"
}

Wan 在运动驱动参考下要求同时提供 start_image 和 end_image。音频未原生暴露,请在后期添加。

全部 12 份模板的适配矩阵

模板 Seedance 2.5 Veo 3.1 Kling 3.0 Wan 3.0 说明
1. 行走人像 @subject image_prompt elements[0] start_image+end_image Wan 需手绘尾帧
2. 产品展示 @subject+转台 image_prompt elements[0] start_image 四者皆可处理
3. 对白节拍 @subject image_prompt elements[0] 适配度弱 Wan 难以捕捉微妙表情
4. 双人对话 @subject+@subject2 两个 image_prompt elements[0..1] 不支持 Kling 胜出
5. 主体+道具 @subject+@prop 两个 image_prompt elements[0..1] 适配度弱 Seedance/Kling 强劲
6. 群戏阵容 @subject ×3 不实用 elements[0..2] 不支持 仅 Kling 一次性可成
7. 电影级调色 @style image_prompt+线索 elements[0] style start_image+线索 四者皆可处理风格
8. 绘画风风格 @style image_prompt+媒介 elements[0] style start_image 四者皆可处理
9. 品牌识别 @style+品牌线索 image_prompt+线索 elements[0] style start_image+线索 四者皆可处理
10. 首尾帧 非原生 非原生 非原生 原生 Wan 胜出
11. 纯镜头运动 @scene+camera image_prompt+camera elements[0]+camera 适配度弱 Veo/Seedance 强劲
12. 动作弧线 @subject+运动 image_prompt+运动 elements[0]+运动 两帧法 四者皆可,需注意细节

7. 姊妹文章与延伸阅读

外部资料:


8. 验证 H3 兼容性

一种可复现的方法,按五组件骨架对 H3 进行打分:

  1. 确认 H3 接受哪些输入。 检查 H3 API 中是否存在参考图像字段(image_prompt、reference_image、init_image、@tag,或 elements[])、主体令牌机制、仅风格支持,以及首尾帧支持。若无法同时确认四项,则将 H3 视为部分已定义。

  2. 从模板 7(电影级调色)开始。 仅风格模式验证最易。无法通过仅风格测试的模型将无法通过其他任何测试。

  3. 运行模板 1(行走人像)。 这是规范的单主体 Ref2V 试验。

  4. 运行模板 4(双人对话)。 这是首个会暴露主体串扰的测试。若 H3 出现面部互换或特征合并,即为明确的失败信号。

  5. 按五个维度为每次输出打分(0 = 未遵循,1 = 部分遵循,2 = 完全遵循):

| 维度 | 检查内容 | |———–|—————| | 锚点保真度 | 主体是否与参考图相似? | | 风格保真度 | (仅风格模式)风格是否匹配? | | 场景遵循度 | 场景是否与文字描述一致? | | 运动遵循度 | 运动是否符合镜头/主体指令? | | 音频遵循度 | 音频是否符合(如支持)? |

10 分中得 8 分及以上,表示 H3 与通用骨架兼容;低于 8 分则意味着 H3 需要自己的适配方案。

  1. 记录失败。 某个组件失败(例如音频)属于功能缺失,而非对 H3 整体的判决。记下失败点,并使用不依赖该组件的模板。

FAQ

什么是参考图生成视频提示模板?

参考图生成视频(Ref2V)提示模板将参考图像与结构化文本搭配,使视频模型生成的输出继承参考图的视觉属性——身份、风格、品牌或运动。规范结构包含五个组件:锚点、主体令牌、场景、运动、音频。请参阅 arXiv:2508.02458。

Ref2V 与 I2V 有何不同?

I2V 动画化单一来源图像并预测合理的运动。Ref2V 将输入视为*参考*——对身份、风格或构图的约束——并与定义新场景的文本配对。当你需要身份锁定或品牌保真时使用 Ref2V;当你只想让静态图像动起来时使用 I2V。

哪个模型最适合多主体参考?

Kling 3.0 通过其带逐主体角色键的 elements 数组实现。Seedance 2.5 通过 @subject 标签支持,是最便捷的内联方案。Veo 3.1 需要使用多个 image_prompt 字段并合成。Wan 3.0 在单次调用中不支持多主体。

我可以将这些模板用于 H3 或 S2V/V2V 吗?

本模板基于有据可查的 API 设计,尚未针对 H3 进行验证;字段名称可能有所不同。五组件骨架应当可以平移,但请按第 8 节对 H3 打分。本模板专门针对 Ref2V——S2V 与 V2V 相关,但语法不可直接复用。

这篇文章适用于 H3 吗?

可以,作为*起点*。我们未对 H3 进行测试。请将模板视为假设,按第 8 节运行测试,并按 H3 的实际 API 接口调整语法。获得独立 H3 数据后我们将更新本文。

Ref2V 提示应该多长?

场景描述请控制在约 120 词以内,主体令牌控制在 2–6 词。过长的描述会重新引入 Ref2V 本应消除的歧义。


结论

2026 年的参考图生成视频提示是一套五组件骨架,可跨 Seedance、Veo、Kling、Wan 复用——一旦经过验证,也可能适用于 H3 及其他即将推出的模型。请将 12 份模板视为起点,按第 8 节为你的输出打分,并根据你所对接的 API 接口调整语法。

Seedance 深入内容,请参阅 Seedance 2.5 提示指南和多参考图 50 素材配套文章。多模态框架相关内容,请参阅多模态参考视频提示。


由 videosprompt.org 编辑团队审校 · 2026 年 10 月


本文局限性

本 12 份模板已基于 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0 的有据可查 API 进行验证。H3 因关键词可发现性需要而出现在标题与 slug 中,我们尚未独立地将 H3 投入这些模板的实测,也缺乏关于其参考图像 API 接口的第一手数据。将其应用于 H3 的读者应将模板视为假设,并按第 8 节操作。H3 上的结果可能与有据可查的模型有所不同。获得独立 H3 数据后我们将修订本文。

分享文章

相关文章

推荐阅读

开始你的下一步

探索更多可能,发现适合你的解决方案。