VideosPrompt VideosPrompt

图生视频角色锁定 Prompt:主体锁定的 I2V 工作流

作者: VideosPrompt 日期: 2026-10-07 14:12:03
图生视频角色锁定 Prompt:主体锁定的 I2V 工作流

TL;DR

  • I2V 与 Ref2V 的区别:I2V 从单张静帧出发,以图像为条件生成运动。Ref2V 从参考照片中提取主体再生成运动。两者都锁定身份,prompt 语法不同。
  • 角色锁定是 prompt 层面的契约:prompt 必须声明保留什么(面容、服饰、配饰、体型)以及允许变化什么(姿态、表情、镜头、场景)。
  • 锁定强度是一段连续谱:软锁定允许服饰、姿态漂移;中锁定保留面容与服装但允许重新摆姿;硬锁定冻结除微表情和物理驱动运动之外的一切。
  • 起始图像质量决定上限:清晰、正面、单主体、姿态中性、高对比度的源图像可大幅降低身份漂移。
  • 各模型锁定机制不同:Seedance 2.5(参考图像+运动视频对)、Veo 3.1(图像 prompt+音频提示)、Kling 3.0(元素面板)、Wan 3.0(首末帧配对)。
  • 后文提供 12 个已验证模板,按软/中/硬/多角色锁定分组,每个置于 text 代码块中。
  • "preserve identity exact"、"no face morphing"、"strict character continuity" 等锁定强度修饰词能显著提升保持度,必须显式声明。

引言:I2V 流水线中”角色锁定”的真正含义

图生视频(I2V)模型接收单张静帧并合成”紧随其后”的运动。大多数 I2V 模型的默认行为是 *创意式插值*:它们会重新打光面部、调整下颌线、把外套换成另一件类似的,或在让运动更”好看”时给角色增龄。角色锁定是一组 prompt 技巧与起始图像选择的组合,告诉模型”不要重新想象这个身份,只让它动起来”。

I2V 中的角色锁定与 Ref2V(参考生视频)中的角色锁定不同。I2V 是以图像为条件的运动:模型看到一帧静帧并延续它,逐像素锁住整个场景的起始状态。Ref2V 是以身份为条件的生成:模型看到参考照片后生成运动,该运动可能与任何单张参考的光照、构图或姿态都不匹配——只锁定身份,而把其余一切交给 prompt。参见关于 Seedance 参考生视频工作流 的配套指南,以及姊妹文章 2026 年 AI 角色一致性 prompt 以了解更多策略。

角色锁定为何重要:

  1. 资产可复用。锁定角色可在系列营销中重复使用——社交广告、讲解视频、培训视频——免去重拍。
  2. 叙事连贯性。连续出镜的主持人、虚拟形象和品牌吉祥物要求观众在多集中识别同一角色。漂移数秒内即破坏沉浸感。
  3. 唇形同步集成。将生成视频与录制配音合成时,静帧与运动之间的面部漂移会暴露为诡异的违和感。

角色锁定是一项可学习、可 prompt 化的技能,但具有概率性:同一 prompt 在 Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0 上表现各异。本文涵盖通用原则、各模型机制以及 12 个即用模板。


I2V 角色锁定全景:各主流模型的处理方式

每款旗舰模型通过不同机制实现角色锁定,但都收敛于同一契约:一张(或一组)静帧定义身份,prompt 定义其余一切,模型负责生成运动。

模型 锁定机制 主要输入 Prompt 作用 锁定上限
Seedance 2.5 参考图像+运动视频对 主体静帧+运动片段 场景、光照、动作、风格 高;多参考可锁定 50 人演员阵容(参见 Seedance 多参考图像 prompt)
Veo 3.1 图像 prompt+音频提示 起始静帧+可选音频 长描述性 prompt 短片表现高;完整回合中较弱
Kling 3.0 元素面板 面部、服饰、身体静帧 简洁 prompt+面板关键词 面部高,配饰中

| 锁定上限补充见各模型章节。以下为 Wan 3.0 的条目:

| Wan 3.0 | 首末帧配对 | 首帧+末帧 | 简短运动 prompt | 端点最高;中段较弱 |

(说明:上表中文翻译为完整译文。如需更紧凑的表格展示,请自行调整格式。)

Seedance 2.5 将参考静帧视为 *身份锚点*,将独立的运动参考视为 *动作源*。prompt 负责场景、光照、风格——运动来自参考视频。


I2V 角色锁定全景图(续)

各模型处理方式详解

模型 锁定机制 主要输入 Prompt 角色 锁定上限
Seedance 2.5 参考图像+运动视频对 主体静帧+运动片段 场景、光照、动作、风格 高;多参考可锁定 50 人演员阵容(参见 Seedance 多参考图像 prompt)
Veo 3.1 图像 prompt+音频提示 起始静帧+可选音频 长描述性 prompt 短片表现高;完整回合中较弱
Kling 3.0 元素面板 面部、服饰、身体静帧 简洁 prompt+面板关键词 面部高,配饰中等
Wan 3.0 首末帧配对 首帧+末帧 简短运动 prompt 端点最高;中段较弱

角色锁定 prompt 公式:六个必含要素

可靠的 I2V 角色锁定 prompt 包含六部分。软锁定可省略一两个,但硬锁定强度下每个要素都不可或缺。

# 要素 作用 示例片段
1 源图像锚点 再次确认静帧为身份源 "Using the provided still as the identity anchor"
2 角色特征描述 命名锁定内容:脸型、发型、体型、肤色、标志性特征 "mid-30s East Asian woman, round face, short black bob, dark brown eyes"
3 场景内容 描述角色行为 "walks through a crowded night market, glancing left and right"
4 镜头 定义构图、焦距、运动 "handheld medium close-up, 35mm, shallow depth of field"
5 风格 编码视觉处理:电影、动画、摄影 "cinematic, anamorphic, slightly desaturated, 2.39:1"
6 一致性检查 告诉模型*不要*做什么 "no face morphing, no wardrobe change, preserve identity exactly"

第一个要素最重要,却也最常被省略。许多 I2V 失败可追溯到一个常见但致命的问题:prompt 把角色描述得精美绝伦,却从未告诉模型静帧不可更改。第 6 要素是叠加锁定强度修饰词的地方。

一个实用的写法(来自 AI 视频 prompt 工程最佳实践)是:身份描述放在源图像锚点 *之后*、动作 *之前*。这与当前视频扩散模型的注意力流向一致——先读身份,再以身份为条件生成运动。


选择合适的起始图像:什么是”优质锁定源”

起始图像是整个锁定的基础。糟糕的源图像无法靠巧妙的 prompt 挽救;优质的源图像即使配平凡 prompt也能成功。优质锁定源的五个属性:

  1. 面部清晰、光照良好。模糊、逆光或被遮挡的面部会在 0.5 秒内产生漂移。眼部对焦,高度 80px 以上。
  2. 正面或四分之三视角。侧面照隐藏一半身份信号。正面或四分之三视角提供最多锚点——双颊、双眉、完整额头。
  3. 单主体,与他人分离。多人画面迫使模型猜测锚点,猜测结果在中段可能改变。
  4. 主体与背景高对比。融入背景会导致轮廓信号弱、边缘伪影多。
  5. 身份特征无遮挡。墨镜、捂嘴、深帽檐阴影——每遮挡一处就丢失一个锚点。

经验法则:如果选角导演能从十个相似配角中一眼挑出此人,该源图像就能很好地锁定。

分辨率不如信号清晰度重要——一张干净的 768px 面部静帧比一张噪点多的 4K 静帧锁定效果更好。中性光照保留肤色保真度;源图像中的暖色或冷色调会带入输出。

对于生成的源(文生图静帧),向下游传递前请确保满足以上五项。


12 个 prompt 模板(按锁定强度组织)

每个模板都是可直接粘贴的完整 prompt。将 [占位符] 替换为你的值。

软锁定(3 个模板)

软锁定允许模型在姿态、光照、服饰上拥有较大创作自由度。用于运动质量比身份忠实度更重要的场景——B-roll、环境镜头、场景过渡空镜。

Template SL-01 — Street B-roll, soft lock, model: universal
Using the provided still as the identity anchor.
Subject: [mid-30s subject, neutral pose, simple background] from the anchor still.
Action: walks away from camera down a tree-lined street, soft natural daylight.
Camera: handheld 50mm, slow push-in, subject centered in frame.
Style: cinematic, naturalistic color, 2.39:1 anamorphic feel.
Consistency: wardrobe may shift subtly with movement; preserve face shape and skin tone; allow natural lighting variation.
Template SL-02 — Ambient scene, soft lock, model: Seedance 2.5 / Veo 3.1
Reference still is the identity anchor. Motion video provides walk cycle.
Subject: [anchor subject, full body, neutral stance] from the still.
Action: enters a sunlit café, places a coffee cup on a table, smiles.
Camera: static medium shot at table height, 35mm, warm tungsten bounce.
Style: editorial lifestyle photography, soft grain, gentle vignette.
Consistency: identity preservation is important but motion realism takes priority; allow minor wardrobe drift toward similar palette; do not change hair color or facial structure.
Template SL-03 — Establishing shot, soft lock, model: Kling 3.0 / Wan 3.0
Use anchor still as subject identity. End frame is the still itself.
Subject: [anchor subject, three-quarter view, single subject frame] from the still.
Action: stands at a window in a high-rise apartment at dusk, city skyline behind, gently turning toward the camera.
Camera: locked-off wide shot, 24mm, deep focus, golden hour palette.
Style: cinematic still photography, slight film bleach, 16:9.
Consistency: preserve identity; minor pose drift toward end frame is acceptable; preserve skin tone and silhouette.

中锁定(3 个模板)

中锁定保留面部与服饰,但允许姿态、手势、视线变化。叙事类内容、讲解类谈话头、社交短视频中最常用的锁定强度。

Template ML-01 — Talking head, medium lock, model: Veo 3.1
Anchor still is the identity source. Subject speaks the line below.
Subject: [anchor subject, sharp face, front view] from the still.
Action: speaks directly to camera with calm, measured delivery, line: "Welcome back to the series — today we look at three prompts that changed how I shoot video."
Camera: locked-off medium close-up, 50mm, eye level, soft key light at 45 degrees.
Style: clean studio look, neutral background, broadcast-safe color.
Consistency: preserve identity exactly; wardrobe must match the still including collar, lapel, and accessories; only the mouth, eyes, and head may move; no face morphing.
Template ML-02 — Walking interview, medium lock, model: Seedance 2.5
Anchor still defines subject identity. Motion reference is the supplied walk-and-talk clip.
Subject: [anchor subject, three-quarter view, walking pace] from the still.
Action: walks along a coastal path while speaking, gesturing occasionally with the right hand.
Camera: tracking alongside at shoulder height, 35mm, shallow depth of field.
Style: documentary cinematography, handheld feel, neutral grade.
Consistency: preserve identity exactly; wardrobe must match including hat, jacket, and bag strap; allow natural hand gestures and head turns; no face morphing; no wardrobe swap.
Template ML-03 — Re-pose sequence, medium lock, model: Kling 3.0
Anchor face and anchor wardrobe stills are the identity sources.
Subject: [anchor subject, two registered elements: face and wardrobe] from the elements panel.
Action: transitions through four poses over six seconds — front, three-quarter left, three-quarter right, and back to front — in a sunlit room.
Camera: slowly orbiting 360 around subject at 1m radius, 50mm, even exposure.
Style: clean editorial, soft shadow, light beige background.
Consistency: preserve identity exactly across all four poses; wardrobe must match in every pose including accessories; only head, neck, and arms may move; no face morphing; strict character continuity.

硬锁定(3 个模板)

硬锁定将主体冻结到小配饰级别——耳环、手表、痣、精确的衣领褶皱。用于系列连续性、品牌内容以及观众已经”见过”该角色的任何制作场景。

Template HL-01 — Series intro shot, hard lock, model: Veo 3.1
Anchor still is the absolute identity source. Subject speaks the line.
Subject: [anchor subject, exact front view, sharp face] from the still.
Action: a single calm breath, slight smile, then delivers the line: "Welcome to episode one."
Camera: locked-off medium close-up, 85mm, eye level, three-point lighting matching the still.
Style: cinematic, polished, brand-grade color.
Consistency: preserve identity exactly including every facial feature, the small mole near the left eyebrow, the earring in the right ear, the exact collar fold of the shirt; only mouth, eyes, and subtle head movement; no face morphing; no wardrobe change; strict character continuity.
Template HL-02 — Action with retained identity, hard lock, model: Seedance 2.5
Anchor still defines subject. Motion reference is a supplied action clip of the same subject.
Subject: [anchor subject, action pose, sharp face] from the still.
Action: climbs three rungs of a fire escape ladder while looking back over the shoulder toward lens.
Camera: low-angle handheld, 35mm, dramatic perspective, dusk light.
Consistency: preserve identity exactly including every facial feature, jacket color, jacket zipper position, watch on left wrist; allow full body motion including legs, arms, and torso; no face morphing; strict character continuity; preserve accessories exactly.
Template HL-03 — Profile with continuity, hard lock, model: Wan 3.0
Start frame: anchor still. End frame: anchor still rotated 90 degrees to subject's left.
Subject: [anchor subject, exact profile view from end frame].
Action: subject turns head slowly from right profile to left profile over four seconds.
Camera: locked-off close-up profile, 100mm, soft window light.
Style: minimalist editorial, neutral backdrop.
Consistency: preserve identity exactly including every facial feature and the visible side of the hair; allow only the rotational motion; no face morphing; strict character continuity; preserve accessories including earring and hair clip.

保持身份的群像(3 个模板,多角色)

多角色锁定是最难的场景,因为模型必须同时锁定两个或更多身份并保持彼此区分。应通过区别性特征而非位置来描述每个主体——当主体交叉时,基于位置的引用会混淆。

Template MC-01 — Two-character dialogue, cast lock, model: Veo 3.1
Two anchor stills define identities: still A (left character), still B (right character).
Subject A: [anchor A, distinguishing features] seated left.
Subject B: [anchor B, distinguishing features] seated right.
Action: A speaks a line, then B responds; both maintain eye contact with each other, not the camera.
Camera: locked two-shot at 50mm, slight low angle, warm key from screen-left.
Style: cinematic dialogue scene, naturalistic color.
Consistency: preserve both identities exactly; do not let A's features drift toward B's or vice versa; preserve each character's wardrobe including all accessories; no face morphing; strict character continuity for both subjects.
Template MC-02 — Three-character group shot, cast lock, model: Seedance 2.5
Three anchor stills define identities. Motion reference is a supplied group action clip.
Subject A: [anchor A, distinctive wardrobe and features], positioned left.
Subject B: [anchor B, distinctive wardrobe and features], positioned center.
Subject C: [anchor C, distinctive wardrobe and features], positioned right.
Action: the trio walks together across an open plaza, B in the lead, C catching up.
Camera: wide tracking shot, 28mm, even daylight, deep focus.
Style: lifestyle commercial, polished, warm grade.
Consistency: preserve all three identities exactly across the entire motion; no face morphing on any subject; wardrobe must match including hats, bags, and visible logos; do not swap wardrobe between subjects; strict character continuity for the full cast.
Template MC-03 — Confrontation scene, cast lock, model: Kling 3.0
Two anchor face elements and two anchor wardrobe elements registered in the elements panel.
Subject A: [anchor A face element + anchor A wardrobe element].
Subject B: [anchor B face element + anchor B wardrobe element].
Action: A and B stand face to face in a narrow corridor; A leans in slightly while speaking.
Camera: handheld over-the-shoulder alternating close-ups, 50mm, cool fluorescent lighting.
Style: dramatic, high-contrast, desaturated.
Consistency: preserve both identities exactly; preserve both wardrobes exactly; do not let A's face drift into B's or vice versa; preserve all accessories including A's pendant and B's wristwatch; no face morphing; strict character continuity for both subjects.

锁定强度修饰词:强化保持度的语言令牌

锁定强度取决于起始图像、prompt 内容以及用于表示约束程度的精确措辞。按升序从软到硬排列,以下修饰词是我们在四款 I2V 模型上验证最可靠的。

修饰词令牌 锁定信号 使用场景
"allow natural lighting variation" 软 B-roll、环境、氛围镜头
"preserve face shape and skin tone" 软-中 生活方式、纪录片、编辑类
"preserve identity" 中 叙事、讲解、社交
"preserve identity exactly" 中-硬 系列连续性、品牌内容
"no face morphing" 硬 面部特征必须保持稳定
"no wardrobe swap" 硬 品牌造型、角色反复出场
"strict character continuity" 硬 多集系列、反复出场的主持人
"preserve accessories exactly" 最硬 手表、耳环、眼镜、痣必须保留
"no identity drift between frames" 最硬 基准/帧级验证

一个常见错误是在软锁定 prompt 中使用过多硬修饰词。当场景需要重新表情(如从微笑过渡到惊讶)时,模型可能将 "no face morphing" + "preserve identity exactly" + "strict character continuity" 解读为冲突。处理情绪弧线时,回退到 "preserve identity"。

另一个常见错误是:不使用任何修饰词,期望模型从语境推断保持意图。模型从动作动词推断运动优先级,而非保持意图。没有显式的保持语言,运动胜出。

试金石:大声朗读 prompt,划出每个描述*角色做什么*的词。如果这些词的数量超过保持令牌的数量,模型会将运动视为目标、将身份视为装饰。重新平衡,直到保持令牌与动作描述权重相当。

Animate Anyone 论文(Hu et al., 2023, arxiv.org/abs/2311.17117)奠定了角色锁定运动合成的基础架构。身份特征与运动特征宜被视为独立信号流,并附以显式的保持约束。


FAQ

I2V 与 Ref2V 中的角色锁定有何区别? 在 I2V 中,模型看到一帧静帧并延续它,不会重新想象主体。在 Ref2V 中,模型看到参考后组合出新的首帧。I2V *逐像素锁定整个起始状态*;Ref2V *仅锁定身份*,释放光照、姿态与构图。

角色锁定在真实照片还是生成图像上效果更好? 相同分辨率下,中性光照下的真实照片锁定效果更好。生成图像常含有微妙扭曲身份的伪影(面部不对称、耳形不一致、耳环几何融化),模型将其视为合法锚点。使用前请对生成静帧进行 QA。

不更换模型,提升角色锁定效果的最经济方法? 改善起始图像。更清晰、更正面、单主体、姿态中性、高对比度的源图像产生的锁定效果强于任何 prompt 微调。

角色锁定能撑过 180 度环绕镜头吗? 勉强可以。大多数模型在大约 90 度弧度内保持身份,之后在后半段漂移,因为头部背面提供的信号更少。360 度拍摄时,应硬锁定可见配饰(发夹、夹克 logo、手表)。Kling 3.0 的元素面板在此场景下处理最佳。

(关于上下文的 I2V 中角色锁定的进一步讨论)


FAQ(续)

如何在多段视频间保持服饰稳定? 将服饰注册为独立锚点——一张清晰的全身服饰静帧或 Kling 3.0 的服饰元素。使用 "preserve wardrobe exactly" 和 "no wardrobe swap" 引用它。如果该造型有标志性配饰,请在 prompt 中点名以建立语言锚点。

锁定强度与运动质量之间是什么关系? 两者此消彼长。硬锁定将运动压扁为僵硬;软锁定优先运动真实感,代价是漂移。品牌安全要硬锁定;MV B-roll 通常要软锁定。在具体片段上两端都测试。

Seedance 2.5、Veo 3.1、Kling 3.0、Wan 3.0 是否有特定的角色锁定限制? Seedance 2.5 在多参考群像上最强(参见 Seedance 多参考图像 prompt)。Veo 3.1 在带音频的单主体叙事上最强。Kling 3.0 在面板注册的面部+服饰主体上最强。Wan 3.0 在端点保真度上最强,在中段漂移上最弱。


结语

I2V 中的角色锁定不是某个功能开关。它是起始图像、锁定强度修饰词与模型保持行为之间的契约。上述 12 个模板只是起点——每个项目都有自己的品牌、选角与漂移容忍度。运行 4 段测试网格(每模型一段,其余一致)以确定其锁定模型对你角色的锁定上限。

更多策略,参见配套文章 2026 年 AI 角色一致性 prompt。Seedance 战术,参见 Seedance 2.5 prompt 指南 和 Seedance 多参考图像 prompt。


由 videosprompt.org 编辑团队审校 · 2026 年 10 月

分享文章

相关文章

推荐阅读

开始你的下一步

探索更多可能,发现适合你的解决方案。