图生视频角色锁定 Prompt:主体锁定的 I2V 工作流
TL;DR
- I2V 与 Ref2V 的区别:I2V 从单张静帧出发,以图像为条件生成运动。Ref2V 从参考照片中提取主体再生成运动。两者都锁定身份,prompt 语法不同。
- 角色锁定是 prompt 层面的契约:prompt 必须声明保留什么(面容、服饰、配饰、体型)以及允许变化什么(姿态、表情、镜头、场景)。
- 锁定强度是一段连续谱:软锁定允许服饰、姿态漂移;中锁定保留面容与服装但允许重新摆姿;硬锁定冻结除微表情和物理驱动运动之外的一切。
- 起始图像质量决定上限:清晰、正面、单主体、姿态中性、高对比度的源图像可大幅降低身份漂移。
- 各模型锁定机制不同:Seedance 2.5(参考图像+运动视频对)、Veo 3.1(图像 prompt+音频提示)、Kling 3.0(元素面板)、Wan 3.0(首末帧配对)。
- 后文提供 12 个已验证模板,按软/中/硬/多角色锁定分组,每个置于
text代码块中。 "preserve identity exact"、"no face morphing"、"strict character continuity"等锁定强度修饰词能显著提升保持度,必须显式声明。
引言:I2V 流水线中”角色锁定”的真正含义
图生视频(I2V)模型接收单张静帧并合成”紧随其后”的运动。大多数 I2V 模型的默认行为是 *创意式插值*:它们会重新打光面部、调整下颌线、把外套换成另一件类似的,或在让运动更”好看”时给角色增龄。角色锁定是一组 prompt 技巧与起始图像选择的组合,告诉模型”不要重新想象这个身份,只让它动起来”。
I2V 中的角色锁定与 Ref2V(参考生视频)中的角色锁定不同。I2V 是以图像为条件的运动:模型看到一帧静帧并延续它,逐像素锁住整个场景的起始状态。Ref2V 是以身份为条件的生成:模型看到参考照片后生成运动,该运动可能与任何单张参考的光照、构图或姿态都不匹配——只锁定身份,而把其余一切交给 prompt。参见关于 Seedance 参考生视频工作流 的配套指南,以及姊妹文章 2026 年 AI 角色一致性 prompt 以了解更多策略。
角色锁定为何重要:
- 资产可复用。锁定角色可在系列营销中重复使用——社交广告、讲解视频、培训视频——免去重拍。
- 叙事连贯性。连续出镜的主持人、虚拟形象和品牌吉祥物要求观众在多集中识别同一角色。漂移数秒内即破坏沉浸感。
- 唇形同步集成。将生成视频与录制配音合成时,静帧与运动之间的面部漂移会暴露为诡异的违和感。
角色锁定是一项可学习、可 prompt 化的技能,但具有概率性:同一 prompt 在 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0 上表现各异。本文涵盖通用原则、各模型机制以及 12 个即用模板。
I2V 角色锁定全景:各主流模型的处理方式
每款旗舰模型通过不同机制实现角色锁定,但都收敛于同一契约:一张(或一组)静帧定义身份,prompt 定义其余一切,模型负责生成运动。
| 模型 | 锁定机制 | 主要输入 | Prompt 作用 | 锁定上限 |
|---|---|---|---|---|
| Seedance 2.5 | 参考图像+运动视频对 | 主体静帧+运动片段 | 场景、光照、动作、风格 | 高;多参考可锁定 50 人演员阵容(参见 Seedance 多参考图像 prompt) |
| Veo 3.1 | 图像 prompt+音频提示 | 起始静帧+可选音频 | 长描述性 prompt | 短片表现高;完整回合中较弱 |
| Kling 3.0 | 元素面板 | 面部、服饰、身体静帧 | 简洁 prompt+面板关键词 | 面部高,配饰中 |
| 锁定上限补充见各模型章节。以下为 Wan 3.0 的条目:
| Wan 3.0 | 首末帧配对 | 首帧+末帧 | 简短运动 prompt | 端点最高;中段较弱 |
(说明:上表中文翻译为完整译文。如需更紧凑的表格展示,请自行调整格式。)
Seedance 2.5 将参考静帧视为 *身份锚点*,将独立的运动参考视为 *动作源*。prompt 负责场景、光照、风格——运动来自参考视频。
I2V 角色锁定全景图(续)
各模型处理方式详解
| 模型 | 锁定机制 | 主要输入 | Prompt 角色 | 锁定上限 |
|---|---|---|---|---|
| Seedance 2.5 | 参考图像+运动视频对 | 主体静帧+运动片段 | 场景、光照、动作、风格 | 高;多参考可锁定 50 人演员阵容(参见 Seedance 多参考图像 prompt) |
| Veo 3.1 | 图像 prompt+音频提示 | 起始静帧+可选音频 | 长描述性 prompt | 短片表现高;完整回合中较弱 |
| Kling 3.0 | 元素面板 | 面部、服饰、身体静帧 | 简洁 prompt+面板关键词 | 面部高,配饰中等 |
| Wan 3.0 | 首末帧配对 | 首帧+末帧 | 简短运动 prompt | 端点最高;中段较弱 |
角色锁定 prompt 公式:六个必含要素
可靠的 I2V 角色锁定 prompt 包含六部分。软锁定可省略一两个,但硬锁定强度下每个要素都不可或缺。
| # | 要素 | 作用 | 示例片段 |
|---|---|---|---|
| 1 | 源图像锚点 | 再次确认静帧为身份源 | "Using the provided still as the identity anchor" |
| 2 | 角色特征描述 | 命名锁定内容:脸型、发型、体型、肤色、标志性特征 | "mid-30s East Asian woman, round face, short black bob, dark brown eyes" |
| 3 | 场景内容 | 描述角色行为 | "walks through a crowded night market, glancing left and right" |
| 4 | 镜头 | 定义构图、焦距、运动 | "handheld medium close-up, 35mm, shallow depth of field" |
| 5 | 风格 | 编码视觉处理:电影、动画、摄影 | "cinematic, anamorphic, slightly desaturated, 2.39:1" |
| 6 | 一致性检查 | 告诉模型*不要*做什么 | "no face morphing, no wardrobe change, preserve identity exactly" |
第一个要素最重要,却也最常被省略。许多 I2V 失败可追溯到一个常见但致命的问题:prompt 把角色描述得精美绝伦,却从未告诉模型静帧不可更改。第 6 要素是叠加锁定强度修饰词的地方。
一个实用的写法(来自 AI 视频 prompt 工程最佳实践)是:身份描述放在源图像锚点 *之后*、动作 *之前*。这与当前视频扩散模型的注意力流向一致——先读身份,再以身份为条件生成运动。
选择合适的起始图像:什么是”优质锁定源”
起始图像是整个锁定的基础。糟糕的源图像无法靠巧妙的 prompt 挽救;优质的源图像即使配平凡 prompt也能成功。优质锁定源的五个属性:
- 面部清晰、光照良好。模糊、逆光或被遮挡的面部会在 0.5 秒内产生漂移。眼部对焦,高度 80px 以上。
- 正面或四分之三视角。侧面照隐藏一半身份信号。正面或四分之三视角提供最多锚点——双颊、双眉、完整额头。
- 单主体,与他人分离。多人画面迫使模型猜测锚点,猜测结果在中段可能改变。
- 主体与背景高对比。融入背景会导致轮廓信号弱、边缘伪影多。
- 身份特征无遮挡。墨镜、捂嘴、深帽檐阴影——每遮挡一处就丢失一个锚点。
经验法则:如果选角导演能从十个相似配角中一眼挑出此人,该源图像就能很好地锁定。
分辨率不如信号清晰度重要——一张干净的 768px 面部静帧比一张噪点多的 4K 静帧锁定效果更好。中性光照保留肤色保真度;源图像中的暖色或冷色调会带入输出。
对于生成的源(文生图静帧),向下游传递前请确保满足以上五项。
12 个 prompt 模板(按锁定强度组织)
每个模板都是可直接粘贴的完整 prompt。将 [占位符] 替换为你的值。
软锁定(3 个模板)
软锁定允许模型在姿态、光照、服饰上拥有较大创作自由度。用于运动质量比身份忠实度更重要的场景——B-roll、环境镜头、场景过渡空镜。
Template SL-01 — Street B-roll, soft lock, model: universal
Using the provided still as the identity anchor.
Subject: [mid-30s subject, neutral pose, simple background] from the anchor still.
Action: walks away from camera down a tree-lined street, soft natural daylight.
Camera: handheld 50mm, slow push-in, subject centered in frame.
Style: cinematic, naturalistic color, 2.39:1 anamorphic feel.
Consistency: wardrobe may shift subtly with movement; preserve face shape and skin tone; allow natural lighting variation.
Template SL-02 — Ambient scene, soft lock, model: Seedance 2.5 / Veo 3.1
Reference still is the identity anchor. Motion video provides walk cycle.
Subject: [anchor subject, full body, neutral stance] from the still.
Action: enters a sunlit café, places a coffee cup on a table, smiles.
Camera: static medium shot at table height, 35mm, warm tungsten bounce.
Style: editorial lifestyle photography, soft grain, gentle vignette.
Consistency: identity preservation is important but motion realism takes priority; allow minor wardrobe drift toward similar palette; do not change hair color or facial structure.
Template SL-03 — Establishing shot, soft lock, model: Kling 3.0 / Wan 3.0
Use anchor still as subject identity. End frame is the still itself.
Subject: [anchor subject, three-quarter view, single subject frame] from the still.
Action: stands at a window in a high-rise apartment at dusk, city skyline behind, gently turning toward the camera.
Camera: locked-off wide shot, 24mm, deep focus, golden hour palette.
Style: cinematic still photography, slight film bleach, 16:9.
Consistency: preserve identity; minor pose drift toward end frame is acceptable; preserve skin tone and silhouette.
中锁定(3 个模板)
中锁定保留面部与服饰,但允许姿态、手势、视线变化。叙事类内容、讲解类谈话头、社交短视频中最常用的锁定强度。
Template ML-01 — Talking head, medium lock, model: Veo 3.1
Anchor still is the identity source. Subject speaks the line below.
Subject: [anchor subject, sharp face, front view] from the still.
Action: speaks directly to camera with calm, measured delivery, line: "Welcome back to the series — today we look at three prompts that changed how I shoot video."
Camera: locked-off medium close-up, 50mm, eye level, soft key light at 45 degrees.
Style: clean studio look, neutral background, broadcast-safe color.
Consistency: preserve identity exactly; wardrobe must match the still including collar, lapel, and accessories; only the mouth, eyes, and head may move; no face morphing.
Template ML-02 — Walking interview, medium lock, model: Seedance 2.5
Anchor still defines subject identity. Motion reference is the supplied walk-and-talk clip.
Subject: [anchor subject, three-quarter view, walking pace] from the still.
Action: walks along a coastal path while speaking, gesturing occasionally with the right hand.
Camera: tracking alongside at shoulder height, 35mm, shallow depth of field.
Style: documentary cinematography, handheld feel, neutral grade.
Consistency: preserve identity exactly; wardrobe must match including hat, jacket, and bag strap; allow natural hand gestures and head turns; no face morphing; no wardrobe swap.
Template ML-03 — Re-pose sequence, medium lock, model: Kling 3.0
Anchor face and anchor wardrobe stills are the identity sources.
Subject: [anchor subject, two registered elements: face and wardrobe] from the elements panel.
Action: transitions through four poses over six seconds — front, three-quarter left, three-quarter right, and back to front — in a sunlit room.
Camera: slowly orbiting 360 around subject at 1m radius, 50mm, even exposure.
Style: clean editorial, soft shadow, light beige background.
Consistency: preserve identity exactly across all four poses; wardrobe must match in every pose including accessories; only head, neck, and arms may move; no face morphing; strict character continuity.
硬锁定(3 个模板)
硬锁定将主体冻结到小配饰级别——耳环、手表、痣、精确的衣领褶皱。用于系列连续性、品牌内容以及观众已经”见过”该角色的任何制作场景。
Template HL-01 — Series intro shot, hard lock, model: Veo 3.1
Anchor still is the absolute identity source. Subject speaks the line.
Subject: [anchor subject, exact front view, sharp face] from the still.
Action: a single calm breath, slight smile, then delivers the line: "Welcome to episode one."
Camera: locked-off medium close-up, 85mm, eye level, three-point lighting matching the still.
Style: cinematic, polished, brand-grade color.
Consistency: preserve identity exactly including every facial feature, the small mole near the left eyebrow, the earring in the right ear, the exact collar fold of the shirt; only mouth, eyes, and subtle head movement; no face morphing; no wardrobe change; strict character continuity.
Template HL-02 — Action with retained identity, hard lock, model: Seedance 2.5
Anchor still defines subject. Motion reference is a supplied action clip of the same subject.
Subject: [anchor subject, action pose, sharp face] from the still.
Action: climbs three rungs of a fire escape ladder while looking back over the shoulder toward lens.
Camera: low-angle handheld, 35mm, dramatic perspective, dusk light.
Consistency: preserve identity exactly including every facial feature, jacket color, jacket zipper position, watch on left wrist; allow full body motion including legs, arms, and torso; no face morphing; strict character continuity; preserve accessories exactly.
Template HL-03 — Profile with continuity, hard lock, model: Wan 3.0
Start frame: anchor still. End frame: anchor still rotated 90 degrees to subject's left.
Subject: [anchor subject, exact profile view from end frame].
Action: subject turns head slowly from right profile to left profile over four seconds.
Camera: locked-off close-up profile, 100mm, soft window light.
Style: minimalist editorial, neutral backdrop.
Consistency: preserve identity exactly including every facial feature and the visible side of the hair; allow only the rotational motion; no face morphing; strict character continuity; preserve accessories including earring and hair clip.
保持身份的群像(3 个模板,多角色)
多角色锁定是最难的场景,因为模型必须同时锁定两个或更多身份并保持彼此区分。应通过区别性特征而非位置来描述每个主体——当主体交叉时,基于位置的引用会混淆。
Template MC-01 — Two-character dialogue, cast lock, model: Veo 3.1
Two anchor stills define identities: still A (left character), still B (right character).
Subject A: [anchor A, distinguishing features] seated left.
Subject B: [anchor B, distinguishing features] seated right.
Action: A speaks a line, then B responds; both maintain eye contact with each other, not the camera.
Camera: locked two-shot at 50mm, slight low angle, warm key from screen-left.
Style: cinematic dialogue scene, naturalistic color.
Consistency: preserve both identities exactly; do not let A's features drift toward B's or vice versa; preserve each character's wardrobe including all accessories; no face morphing; strict character continuity for both subjects.
Template MC-02 — Three-character group shot, cast lock, model: Seedance 2.5
Three anchor stills define identities. Motion reference is a supplied group action clip.
Subject A: [anchor A, distinctive wardrobe and features], positioned left.
Subject B: [anchor B, distinctive wardrobe and features], positioned center.
Subject C: [anchor C, distinctive wardrobe and features], positioned right.
Action: the trio walks together across an open plaza, B in the lead, C catching up.
Camera: wide tracking shot, 28mm, even daylight, deep focus.
Style: lifestyle commercial, polished, warm grade.
Consistency: preserve all three identities exactly across the entire motion; no face morphing on any subject; wardrobe must match including hats, bags, and visible logos; do not swap wardrobe between subjects; strict character continuity for the full cast.
Template MC-03 — Confrontation scene, cast lock, model: Kling 3.0
Two anchor face elements and two anchor wardrobe elements registered in the elements panel.
Subject A: [anchor A face element + anchor A wardrobe element].
Subject B: [anchor B face element + anchor B wardrobe element].
Action: A and B stand face to face in a narrow corridor; A leans in slightly while speaking.
Camera: handheld over-the-shoulder alternating close-ups, 50mm, cool fluorescent lighting.
Style: dramatic, high-contrast, desaturated.
Consistency: preserve both identities exactly; preserve both wardrobes exactly; do not let A's face drift into B's or vice versa; preserve all accessories including A's pendant and B's wristwatch; no face morphing; strict character continuity for both subjects.
锁定强度修饰词:强化保持度的语言令牌
锁定强度取决于起始图像、prompt 内容以及用于表示约束程度的精确措辞。按升序从软到硬排列,以下修饰词是我们在四款 I2V 模型上验证最可靠的。
| 修饰词令牌 | 锁定信号 | 使用场景 |
|---|---|---|
"allow natural lighting variation" |
软 | B-roll、环境、氛围镜头 |
"preserve face shape and skin tone" |
软-中 | 生活方式、纪录片、编辑类 |
"preserve identity" |
中 | 叙事、讲解、社交 |
"preserve identity exactly" |
中-硬 | 系列连续性、品牌内容 |
"no face morphing" |
硬 | 面部特征必须保持稳定 |
"no wardrobe swap" |
硬 | 品牌造型、角色反复出场 |
"strict character continuity" |
硬 | 多集系列、反复出场的主持人 |
"preserve accessories exactly" |
最硬 | 手表、耳环、眼镜、痣必须保留 |
"no identity drift between frames" |
最硬 | 基准/帧级验证 |
一个常见错误是在软锁定 prompt 中使用过多硬修饰词。当场景需要重新表情(如从微笑过渡到惊讶)时,模型可能将 "no face morphing" + "preserve identity exactly" + "strict character continuity" 解读为冲突。处理情绪弧线时,回退到 "preserve identity"。
另一个常见错误是:不使用任何修饰词,期望模型从语境推断保持意图。模型从动作动词推断运动优先级,而非保持意图。没有显式的保持语言,运动胜出。
试金石:大声朗读 prompt,划出每个描述*角色做什么*的词。如果这些词的数量超过保持令牌的数量,模型会将运动视为目标、将身份视为装饰。重新平衡,直到保持令牌与动作描述权重相当。
Animate Anyone 论文(Hu et al., 2023, arxiv.org/abs/2311.17117)奠定了角色锁定运动合成的基础架构。身份特征与运动特征宜被视为独立信号流,并附以显式的保持约束。
FAQ
I2V 与 Ref2V 中的角色锁定有何区别? 在 I2V 中,模型看到一帧静帧并延续它,不会重新想象主体。在 Ref2V 中,模型看到参考后组合出新的首帧。I2V *逐像素锁定整个起始状态*;Ref2V *仅锁定身份*,释放光照、姿态与构图。
角色锁定在真实照片还是生成图像上效果更好? 相同分辨率下,中性光照下的真实照片锁定效果更好。生成图像常含有微妙扭曲身份的伪影(面部不对称、耳形不一致、耳环几何融化),模型将其视为合法锚点。使用前请对生成静帧进行 QA。
不更换模型,提升角色锁定效果的最经济方法? 改善起始图像。更清晰、更正面、单主体、姿态中性、高对比度的源图像产生的锁定效果强于任何 prompt 微调。
角色锁定能撑过 180 度环绕镜头吗? 勉强可以。大多数模型在大约 90 度弧度内保持身份,之后在后半段漂移,因为头部背面提供的信号更少。360 度拍摄时,应硬锁定可见配饰(发夹、夹克 logo、手表)。Kling 3.0 的元素面板在此场景下处理最佳。
(关于上下文的 I2V 中角色锁定的进一步讨论)
FAQ(续)
如何在多段视频间保持服饰稳定?
将服饰注册为独立锚点——一张清晰的全身服饰静帧或 Kling 3.0 的服饰元素。使用 "preserve wardrobe exactly" 和 "no wardrobe swap" 引用它。如果该造型有标志性配饰,请在 prompt 中点名以建立语言锚点。
锁定强度与运动质量之间是什么关系? 两者此消彼长。硬锁定将运动压扁为僵硬;软锁定优先运动真实感,代价是漂移。品牌安全要硬锁定;MV B-roll 通常要软锁定。在具体片段上两端都测试。
Seedance 2.5、Veo 3.1、Kling 3.0、Wan 3.0 是否有特定的角色锁定限制? Seedance 2.5 在多参考群像上最强(参见 Seedance 多参考图像 prompt)。Veo 3.1 在带音频的单主体叙事上最强。Kling 3.0 在面板注册的面部+服饰主体上最强。Wan 3.0 在端点保真度上最强,在中段漂移上最弱。
结语
I2V 中的角色锁定不是某个功能开关。它是起始图像、锁定强度修饰词与模型保持行为之间的契约。上述 12 个模板只是起点——每个项目都有自己的品牌、选角与漂移容忍度。运行 4 段测试网格(每模型一段,其余一致)以确定其锁定模型对你角色的锁定上限。
更多策略,参见配套文章 2026 年 AI 角色一致性 prompt。Seedance 战术,参见 Seedance 2.5 prompt 指南 和 Seedance 多参考图像 prompt。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章