多镜头分镜 AI 视频提示词:从镜头清单到单次生成
要点速览
- 多镜头分镜 AI 视频提示词把整份镜头清单——若干切点、机位设置与节拍——压缩进一次生成请求,而不是手动拼接片段。
- 2026 年通行的格式是带时间戳的:逐秒节拍(”0:00–0:04 — SHOT 1: wide…”)给模型一条它真正能遵循的可编辑时间线。
- 转场词汇的支持程度参差不齐:硬切和叠化被大多数前沿模型原生处理;甩摇、匹配剪辑和 J 剪辑则经常被幻觉处理或直接忽略。
- 连贯性锚点——固定的角色块,加上”所有镜头中 X 相同”的条款以及环境延续——才是让主体跨镜头边界保持稳定的关键。
- 机位指令必须在每个镜头边界处重置,否则模型会把上一镜头的运动延续进下一镜头。
- 本指南包含 12 个可直接套用的提示词模板,覆盖四类故事:产品揭示、角色叙事、教程和节拍驱动的蒙太奇。
- 当你需要超过约 30 秒的时长、逐镜头重拍或精确到帧的时序时,应拆分为链式生成;Seedance 2.5 的时间戳编辑覆盖了介于两者之间的地带。
2026 年的多镜头生成:变化何在
在 AI 视频的大部分历程中,”多镜头”意味着多次生成。你为每个镜头写一条提示词,生成五个片段,然后把整个下午耗在剪辑软件里,设法让同一个角色在所有片段中都穿着同一件夹克。分镜是存在的——只不过它活在模型之外。
这已不再是默认工作流。2026 年的三项进展重塑了分镜变成视频的方式:
Seedance 2.5 带来了时间戳控制与多轮扩展。 你可以指定时间线上确切位置的事件(”at 0:08, the camera arrives at the window”),并通过多轮扩展把片段延伸到约 30 秒的叙事素材,按时间戳编辑,而不是整体重新生成。中新社 2026 年 8 月的报道介绍了 Seedance 2.5 在 API 层面的多镜头叙事能力,把该模型的 API 定位为叙事工具而非片段工厂(chinanews.com.cn)。
Kling 3.0 加入了原生多镜头支持——单条提示词最多 6 个镜头。 不再是让模型自行想象一个场景并指望它切镜,而是把提示词组织成六个离散的镜头块,模型会在一次生成中把它们渲染为六组离散的机位设置(mstudio.ai)。
时间戳语言成了一项通用的提示词技能。 即使在没有正式时间戳 API 的模型上,逐秒的节拍结构如今也能可验证地改善镜头边界表现,因为它迫使提示词去描述一条时间线,而不是一种氛围。
为什么分镜依然重要——事实上对 AI 视频而言反而*更加*重要:生成模型对你的叙事弧毫无概念,除非你把它编码进去。镜头清单是编码叙事意图最古老的接口:它指明观众看到什么、以何种顺序、看多久,以及画面之间如何衔接。多镜头分镜 AI 视频提示词,不过是这份文档被翻译成了模型理解的语法而已。
本指南的余下部分,就是这本翻译手册。
分镜到提示词的翻译
拿一份传统镜头清单来说。这是导演在商业拍摄前交给摄影指导的那种文档:
SHOT 1: WIDE — A ceramic pour-over kettle on a concrete counter, steam rising. Morning light from camera left.
SHOT 2: CU — Water hitting coffee grounds, bloom forming. Shallow focus.
SHOT 3: TILT DOWN — The finished cup placed on a linen mat, logo visible.
三个镜头。三组机位。每两者之间一个切点。人类剧组读完就知道该怎么做。而视频模型读到的却是”一个布光如电影般漂亮的精美手冲咖啡场景”这样一段文字,然后还给你一个连续的、毫无调度的、永不切镜的长镜头。
翻译包含三个动作:
动作 1 — 给镜头编号,并为每个镜头分配时长。 镜头清单的隐含结构由此变为显式。多镜头提示词中的时长就是预算分配:如果模型总共支持 10 秒,4/3/3 的分配方式会告诉它重心放在哪里。如果你不分配时长,模型就会替你分配——通常还分配错了,把一半的片段花在定场镜头上。
动作 2 — 把机位术语转换成带方向的句子。 “WIDE” 变成 “static wide shot, camera locked off”;”TILT DOWN” 变成 “the camera tilts downward from… to…”。模型对”有起点也有终点的句子式运动描述”有反应,对景别缩写则没有。
动作 3 — 把它放到时间线上。 这是承重的一步。带时间戳的结构如下所示:
Template — Storyboard-to-prompt translation (generic 3-shot)
0:00–0:04 — SHOT 1: static wide shot. A ceramic pour-over kettle on a concrete counter, steam rising, morning light from camera left. Camera locked off, no movement.
CUT to SHOT 2.
0:04–0:07 — SHOT 2: macro close-up. Hot water hits fresh coffee grounds, the bloom expands. Shallow depth of field, camera pushes in slowly.
CUT to SHOT 3.
0:07–0:10 — SHOT 3: medium shot. The finished cup is placed on a linen mat, the logo on the cup facing camera. The camera tilts down to rest on the cup and holds static.
注意这个结构做了什么:每一行都以时间范围开头,每个镜头都有一个硬编号(”SHOT 2”),转场位于范围之间的边界而不是镜头内部,而且每个镜头都以一个明确的机位结束状态(”holds static”)收尾。模型不再是在解读一个场景;它是在填充一条你已经设计好的时间线。
这也是为什么,即使在没有时间戳 API 的平台上,时间戳依然重要。这种结构强制了完整性——一份没有缺口、没有无动机时刻、没有镜头彼此意外融混的镜头清单。
镜头转场语言:模型听到的,与你想表达的
转场是分镜到提示词契约中对称性最差的部分。有些转场词汇实际上是原生的——模型见过足够多的剪辑素材,能正确渲染切点;有些被部分理解;还有些会被默默忽略,或幻觉成你根本没有要求的东西。
以下是我们整理的、截至 2026 年 10 月前沿模型处理转场词汇的工作图谱:
| 转场 | 有效的提示词措辞 | 支持程度 | 不受支持时的失败模式 |
|---|---|---|---|
| 硬切 | “CUT TO SHOT 2” / “SMASH CUT” | 原生——默认行为 | 无;如果没出现切点,说明提示词从未要求切点 |
| 叠化 / 交叉溶解 | “the image dissolves into shot 2” | 大多数前沿模型原生支持 | 变成缓慢的推轨移动,而不是画面融合 |
| 淡入黑场 / 由黑场淡出 | “fades to black, then fades up on…” | 基本原生 | 时点漂移;淡入淡出吃掉 2 秒以上的预算 |
| 匹配剪辑 | “CUT: the round coffee cup matches into the round wheel of the bicycle” | 部分——两帧都有描述时才有效 | 模型切了却没做匹配;图形上的呼应丢失 |
| 甩摇 | “the camera whips right, blurring into shot 2” | 部分——Seedance 和 Kling 处理方向性甩摇最好 | 被幻觉成一次切点,或甩摇把整帧涂抹开 |
| J 剪 / L 剪(声音先行) | “audio of shot 2 begins under shot 1, then cut” | 差——大多数模型忽略声音预叠 | 被默默丢弃;视觉切点照常发生 |
| 急推变焦转场 | “snap zoom into the subject, cut on the zoom” | 部分 | 变焦发生了但没有切点,或切了但没有变焦 |
| 圆形划像 / 擦除 / 翻页 | “iris wipe to shot 2” | 在照片级真实模型上差,在风格化模型上更好 | 被替换成叠化,或被完全忽略 |
由此得出两条实用规则。
规则 1:每次都把硬切写进文字。 因为”切”是模型的原生转场,缺少明确的切点指令往往就等于没有切点——模型会给你一个长镜头。多镜头提示词中的每个镜头边界都应当带上明确的 “CUT to SHOT n”。仅凭这一个习惯,就能比其他任何办法修复更多失败的多镜头生成。
规则 2:用空间方式描述转场,而不只是叫出名字。 单独一句”匹配剪辑”只是一个模型可能遵守、也可能不遵守的标签。”圆形的杯沿切到圆形的方向盘,同样大小,同样位于画面中央”描述的才是机制,而模型渲染机制的可靠程度远高于渲染剪辑术语。甩摇同理:说出方向、出画的帧和入画的帧。
关于把转场机制作为独立技术的更深入探讨——包括以产品为重点的变形转场——请参阅我们的产品转场效果视频提示词指南。
跨镜头的连贯性锚点
多镜头提示词能立刻解决切镜问题,却完全解决不了连贯性问题——除非你把连贯性设计进去。放任不管的话,一个渲染六个镜头的模型,会渲染出六个略有不同的角色,身处六个略有不同的房间里。
多镜头提示词中的连贯性依赖三个机制。
1. 持久角色块
用一个固定块把主体定义一次,然后在主体出现的每个镜头中逐字重复——不是”差不多”重复,而是逐字重复。一个紧凑的块如下所示:
CHARACTER (same in all shots): a 30-34 year old woman, shoulder-length auburn hair with a side part, green eyes, small beauty mark above the left eyebrow, athletic petite build; wearing a navy wool blazer over a cream silk blouse, thin gold chain necklace, no earrings.
“same in all shots” 这个短语不是装饰。它告诉模型:这段描述是一条全局约束,而不是一个可以重新随机的单镜头细节。同样的模式也适用于非人类主体:”same product in all shots: matte-black cylindrical smartwatch with a cream face and a single gold crown.”
2. 环境延续条款
角色只是连贯性的一半。房间同样会漂移:混凝土台面变成大理石,窗户移到对面的墙上,晨光变成黄昏。用一个重申常量的环境块来对抗这一点:
ENVIRONMENT (continuous across all shots): a minimal kitchen with polished concrete counters, one large window camera-left, warm morning sunlight, a linen mat beside the sink; same set, same lighting, same time of day in every shot.
陈述不变量——”same set, same lighting, same time of day”——而不是每个镜头都从零重新描述房间。重新描述会招致重新解读;不变量陈述则邀请模型保留原样。
3. 跨镜头锚点引用
在每个镜头块内部,引用锚点而不是重新发明它们:”the woman (character block as defined) sets down the cup.” 这让提示词更短,也在结构上清楚表明所有镜头都取自同一个身份。
在支持参考图像的模型上,文本锚点还可以与视觉锚点配对。我们关于用参考图生视频工作流保持分镜连贯性的姊妹文章介绍了如何在多镜头提示词下堆叠参考帧——当镜头较长时,这是最能保持身份一致的组合。
完整的角色一致性工具包——七词元锚点、漂移分类和验证清单——收录在我们的姊妹指南Kling 3.0 角色一致性中,同样的锚点格式可以直接套进本文的每一个模板。
逐镜头的机位词汇:在每个边界处重置
下面是每个初次撰写多镜头提示词的人都会撞上的失败模式:镜头 1 以 “the camera dollies in slowly” 开场,之后的每个镜头也都在推轨——因为模型把摄影机运动当成片段的全局属性,而不是某个镜头的局部属性。到镜头 4 时,你要求的”static close-up”仍在缓缓向前爬。
修复方法是在每个镜头边界处重置。三部分纪律:
- 每个镜头块都以明确的摄影机状态开头。 即便是 “camera static” 也好。永远不要假定模型会延续你的*意图*——要假定它延续的是*上一段运动*,并把它覆盖掉。
- 说出动作、轴向和结束状态。 “The camera arcs 30 degrees to the right around the cup and settles static” 给出了起点、路径和终点;”Cinematic camera movement” 什么都没给。
- 最后一个镜头以驻留收尾。 最后一个镜头应写明 “camera holds static”,让片段收住,而不是到最后一帧还在移动——如果你打算在第二轮扩展这个片段,这一点就很关键。
一个能胜任这件事的镜头块头部:
0:04–0:08 — SHOT 2: medium close-up of the woman at the window. CAMERA: new setup — static, eye-level, 50mm framing. She turns toward the light and smiles. Camera does NOT continue any previous movement.
括号里的 “new setup” 和否定指令(”does NOT continue any previous movement”)都有帮助:模型对明确的重置有响应,而否定式机位指令是少数能够可靠生效的否定之一——因为不存在把它推向相反方向的竞争性视觉先验。
更广阔的词汇——甩摇、急推变焦、弧形运镜、无人机升起,以及哪些模型能忠实渲染哪些动作——收录在我们的电影级运镜提示词指南中。可以从那里取用单个动作;这里的规则是,每个动作只属于一个镜头。至于底层的电影摄影术语本身——镜头语言、构图术语,以及职业电影人实际使用的运动词汇——Google AI Studio Veo prompt cookbook 上的商业提示词合集是可靠参考,展示了专业人士在提示词到达模型之前如何措辞机位指令。
按故事类型划分的 12 个提示词模板
下面每个模板都是一份完整的、可直接编辑的多镜头分镜 AI 视频提示词。时长按总长约 10–15 秒设定;更长的格式请按比例缩放。角色块和环境块都以内联方式书写,方便你整体替换而无需重构结构。
产品揭示弧(3 个镜头 → 1 条提示词)
模板 1 — 悬念 → 揭示 → 主体驻留
Template — Product reveal: tease / reveal / hero (3 shots, ~12s)
SAME PRODUCT IN ALL SHOTS: matte-black cylindrical smartwatch with a cream watch face and one gold crown, resting on its brushed-metal charging puck.
ENVIRONMENT (continuous across all shots): dark walnut table, single warm key light camera-right, deep black seamless background, faint haze in the air.
0:00–0:04 — SHOT 1: extreme close-up, macro. Only the watch crown and a sliver of the black case are visible, rim-lit by the key light; the rest falls into shadow. CAMERA: static, locked off on a macro lens.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: the camera pulls back smoothly to a medium shot as the key light widens, revealing the full watch face, the cream dial catching the light. CAMERA: new setup — slow straight pull-back, ending static with the watch centered.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: hero product shot, eye-level medium close-up. The watch rotates 90 degrees on its puck into a perfect profile, then holds. CAMERA: static throughout, shallow depth of field. Camera holds static on the final frame.
模板 2 — 细节级联(逐一呈现特性)
Template — Product reveal: detail cascade (3 shots, ~12s)
SAME PRODUCT IN ALL SHOTS: a slim silver smartphone with a matte glass back and a triple camera array, no logos.
ENVIRONMENT (continuous across all shots): pale gray studio sweep, soft overhead diffuser, gentle gradient shadow beneath the product.
0:00–0:04 — SHOT 1: close-up of the phone lying face-down; light rakes across the matte back, revealing the texture. CAMERA: static, high angle looking straight down.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: macro close-up of the triple camera array; a highlight sweeps across the lens rings. CAMERA: new setup — slow lateral truck left to right, ending static.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: the phone lifts upright into a floating three-quarter hero view, screen off, slowly turning 45 degrees. CAMERA: new setup — static, eye-level. Camera holds static.
模板 3 — 问题 → 转变 → 成果
Template — Product reveal: problem / transformation / payoff (3 shots, ~12s)
SAME PRODUCT IN ALL SHOTS: a matte-white wireless earbud in its pebble-shaped charging case.
ENVIRONMENT (continuous across all shots): warm oak desk by a window, soft daylight camera-left, coffee ring on the desk as a fixed landmark.
0:00–0:04 — SHOT 1: wide shot of the cluttered desk; tangled wired earphones sit center frame, the white case pushed to the edge. CAMERA: static, slightly high angle.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: close-up — a hand sweeps the wired earphones out of frame and slides the pebble case to center; the case lid pops open. CAMERA: new setup — static, top-down.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: macro shot inside the open case, both earbuds seated, a small LED breathing once. CAMERA: slow push-in, ending static on the LED. Camera holds static.
角色叙事(3 个)
模板 4 — 出发节拍
Template — Character narrative: departure (3 shots, ~12s)
CHARACTER (same in all shots): a 30-34 year old woman, shoulder-length auburn hair with a side part, green eyes, small beauty mark above the left eyebrow, petite athletic build; wearing a navy wool blazer over a cream silk blouse, thin gold chain necklace.
ENVIRONMENT (continuous across all shots): a quiet apartment hallway with white walls, wooden floor, one window at the far end, cool overcast morning light; same set and lighting in every shot.
0:00–0:04 — SHOT 1: medium shot. The woman stands at the hallway end, keys in hand, looking back toward the door once. CAMERA: static, eye-level.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: close-up of her hand turning the brass door handle. CAMERA: new setup — static, tight framing on the hand.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: wide shot from inside — she steps out through the door into the light from the far window and the door swings shut. CAMERA: new setup — static, camera behind her. Camera holds static.
模板 5 — 顿悟节拍
Template — Character narrative: realization (3 shots, ~12s)
CHARACTER (same in all shots): a 25-29 year old man with short curly black hair, wire-rim glasses, warm brown skin, slim build; wearing a charcoal crew-neck sweater, a canvas messenger bag strap across his chest.
ENVIRONMENT (continuous across all shots): a city bus stop at dusk, wet pavement reflecting streetlights, blurred traffic behind; same location, same dusk light in every shot.
0:00–0:04 — SHOT 1: medium shot. He checks his empty wrist — no watch — then pats the messenger bag, face tightening. CAMERA: static, eye-level, 50mm.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: close-up, his face; his eyes widen slightly as he remembers. CAMERA: new setup — very slow push-in, ending static on his eyes.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: wide shot — he turns and runs back the way he came, out of frame left, leaving the empty bus stop. CAMERA: new setup — static. Camera holds static on the empty stop.
模板 6 — 双人对手戏
Template — Character narrative: exchange (3 shots, ~12s)
CHARACTER A (same in all shots): a 30-34 year old woman, shoulder-length auburn hair, green eyes, navy wool blazer over cream silk blouse, gold chain necklace.
CHARACTER B (same in all shots): a 40-44 year old man, close-cropped gray hair, trimmed beard, navy peacoat over a charcoal scarf.
ENVIRONMENT (continuous across all shots): a café interior with brass fixtures, marble counter, large window camera-right, warm tungsten light.
0:00–0:04 — SHOT 1: over-the-shoulder from behind the man — the woman slides a small envelope across the marble counter, steady eye contact. CAMERA: static.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: reverse over-the-shoulder from behind the woman — the man looks down at the envelope, then back up, and gives one slow nod. CAMERA: new setup — static reverse angle.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: wide two-shot, both in frame at the counter; the man takes the envelope and the woman steps back, releasing the frame. CAMERA: static, eye-level. Camera holds static.
教程 / 操作步骤序列(3 个)
模板 7 — 三步流程
Template — Tutorial: three-step process (3 shots, ~15s)
HANDS (same in all shots): adult hands, short unpolished nails, plain silver band on the right ring finger.
ENVIRONMENT (continuous across all shots): white marble kitchen counter, even soft daylight, overhead camera position; same set and lighting in every shot.
0:00–0:05 — SHOT 1: top-down. Hands place a glass mixing bowl on the counter and crack two eggs into it. CAMERA: static, straight overhead.
CUT to SHOT 2.
0:05–0:10 — SHOT 2: top-down. Hands whisk the eggs in tight circles until uniform pale yellow. CAMERA: new setup — static overhead, slightly tighter framing.
CUT to SHOT 3.
0:10–0:15 — SHOT 3: 45-degree angle. The mixture pours from the bowl into a heated pan, spreading into a perfect circle. CAMERA: new setup — static. Camera holds static as the circle sets.
模板 8 — 界面 / 屏幕演示
Template — Tutorial: screen walkthrough (3 shots, ~12s)
SCREEN (continuous across all shots): the same minimal analytics dashboard, dark theme, one line chart in teal, no readable text — abstract UI only.
ENVIRONMENT (continuous across all shots): a matte-black laptop on a light oak desk, soft window light camera-left.
0:00–0:04 — SHOT 1: wide shot — the laptop open on the desk, dashboard visible, a hand resting beside the trackpad. CAMERA: static, eye-level.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: close-up of the screen — the cursor clicks a highlighted card, which expands with a smooth panel animation. CAMERA: new setup — static, straight on, no glare.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: the teal line on the chart draws upward in one continuous stroke; the panel glows faintly at the peak. CAMERA: slow push-in, ending static on the peak. Camera holds static.
模板 9 — 前后对比转变
Template — Tutorial: before / after (3 shots, ~12s)
OBJECT (same in all shots): a scuffed brown leather boot pair with worn laces.
ENVIRONMENT (continuous across all shots): a walnut workbench under a single overhead lamp, brush and cloth laid out in fixed positions.
0:00–0:04 — SHOT 1: top-down static shot of the scuffed boots side by side, clearly worn. CAMERA: static.
CUT to SHOT 2.
0:04–0:08 — SHOT 2: close-up — a cloth works conditioner into the leather in slow circular strokes; the scuff visibly darkens and evens. CAMERA: new setup — static, tight on the stroke.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: identical framing to SHOT 1 — the boots now restored, a deep even shine, laces crisply tied. CAMERA: static, straight down, matching SHOT 1 exactly. Camera holds static.
音乐 / 节拍驱动蒙太奇(3 个)
模板 10 — 城市节奏蒙太奇
Template — Montage: city rhythm (3 shots, ~9s — one shot per beat, ~3s each)
STYLE (all shots): high-contrast night grade, teal-and-orange practicals, light film grain.
NO characters required; incidental silhouettes only.
0:00–0:03 — SHOT 1: low angle — sneakers crossing a wet crosswalk, reflections streaking underfoot. CAMERA: static, ankle height.
CUT (hard cut on the beat) to SHOT 2.
0:03–0:06 — SHOT 2: a neon sign flickers on above a doorway, buzzing to life. CAMERA: static, slight low angle.
HARD CUT on the beat to SHOT 3.
0:06–0:09 — SHOT 3: a subway train blurs past the platform edge in a streak of light, stopping hard. CAMERA: static on the platform. Camera holds static as the train stops.
模板 11 — 产品踩点蒙太奇
Template — Montage: product on beat (3 shots, ~9s)
SAME PRODUCT IN ALL SHOTS: a coral-red sneaker with a chunky white sole.
ENVIRONMENT changes are intentional (montage locations), but the product never changes; same coral-red colorway, same sole, same lacing in every shot.
0:00–0:03 — SHOT 1: the sneaker drops onto concrete in slow motion, sole compressing on impact. CAMERA: static, ground level.
HARD CUT to SHOT 2.
0:03–0:06 — SHOT 2: side profile — the shoe mid-stride against a plain sunlit wall, shadows sharp. CAMERA: tracking laterally at fixed distance, speed matched to the stride.
HARD CUT to SHOT 3.
0:06–0:09 — SHOT 3: the sneaker lands center frame on a studio sweep, bouncing once and settling. CAMERA: static, eye-level of the shoe. Camera holds static.
模板 12 — 情绪渐强蒙太奇
Template — Montage: crescendo (3 shots, ~12s)
CHARACTER (same in all shots): a 30-34 year old woman, shoulder-length auburn hair, green eyes, navy wool blazer, gold chain necklace.
ENVIRONMENT (continuous across all shots): an empty concert hall with rows of red seats, one shaft of stage light; same hall, same light in every shot.
0:00–0:04 — SHOT 1: wide — she sits alone center row, small in the vast hall, head low. CAMERA: static, high angle from the balcony.
DISSOLVE to SHOT 2.
0:04–0:08 — SHOT 2: medium — she lifts her head toward the stage light, a slow breath. CAMERA: new setup — slow push-in, ending static at medium framing.
CUT to SHOT 3.
0:08–0:12 — SHOT 3: close-up — she stands, moving up into the light, which blooms around her shoulders. CAMERA: tilt up with her, settling static on her face. Camera holds static.
这些模板可以组合使用。一条带角色节拍的产品揭示,就是模板 1 的结构再铆接上模板 4 的角色块——保持时间戳连续,逐镜头重置摄影机,并重复锚点即可。
何时拆分为多次生成
单次生成的多镜头是默认选项,而不是标准答案。请使用这张决策表:
| 场景 | 单次生成 | 链式生成 | 时间戳编辑(Seedance 2.5 式) |
|---|---|---|---|
| 2–6 个镜头,总长 ≤ 约 15 秒 | ✅ 最佳匹配——一条提示词,一个连贯性场 | 过度;给接缝处增加漂移风险 | 没必要 |
| 6 个镜头,≤ 约 15 秒(Kling 3.0 原生) | ✅ Kling 的原生镜头数正好覆盖 | 拆分反而违背模型的强项 | — |
| 15–30 秒、节拍清晰的叙事 | ✅ 若模型支持扩展(Seedance 多轮) | 扩展出现漂移时的退路 | ✅ 按时间戳编辑单个节拍,而不是重新随机 |
| 总时长超过 30 秒 | ❌ 超出单次生成的实用长度 | ✅ 链接片段;把锚点带进每一条提示词 | 部分可行——在每个分段内部编辑 |
| 某个镜头需要重拍 | ❌ 重新生成会把一切重新随机 | ✅ 只重新生成失败的镜头,再在剪辑器中重新拼接 | ✅ 最佳选项——重新定位带时间戳的范围 |
| 精确到帧的时长(口播、音画同步) | 不可靠——镜头内时序由模型自行分配 | ✅ 逐镜头提示词提供帧级控制 | ⚠️ 接近,但需验证;并非帧级精确 |
| 每个场景地点/服装不同 | 镜头较少时可用单条提示词 | ✅ 按场景拆分——每个场景一条提示词 | 使用按场景的分段 |
| 极高的机位复杂度(每个镜头多于 1 个运动) | 镜头之间存在运动串扰的风险 | ✅ 每次生成只放一个复杂运动 | 支持的话按镜头编辑 |
三条不值得为它们单开一张表的经验法则:
- 连贯性场会随长度变薄。 每增加一个镜头,角色块就多一次漂移的机会。如果镜头 5 比镜头 1–3 更重要,考虑拆分,让重要的镜头独享一整次生成的模型注意力。
- 用锚点来链接,而不是用希望。 链式生成应以与原提示词完全相同的角色块和环境块开头——外加一行写明前一片段的结束状态(”picking up directly where the previous clip ended: she is standing, facing camera left”)。链式生成总是在接缝处失败;把接缝描述出来。
- 当只有时长不对时,时间戳编辑胜过重新生成。 如果素材是对的,只是镜头 2 超时了,一次针对时间戳的编辑能保留其余一切。为了修两秒钟而把整条提示词重新随机一遍,就是连贯性死掉的方式。
常见问题
什么是多镜头分镜 AI 视频提示词? 它是一条编码了整份镜头清单的单一提示词——多个镜头、明确的切点、逐镜头的机位指令以及逐秒的时序——使一次生成产出的是一段剪辑好的序列,而不是一个连续长镜头。它取代了过去那种每个镜头生成一个片段、再手动拼接的工作流。
哪些模型在单次生成中支持最多的镜头? 截至 2026 年 10 月,Kling 3.0 在单条提示词中原生支持最多 6 个镜头(mstudio.ai),而 Seedance 2.5 侧重时间戳控制和多轮扩展以实现叙事长度(chinanews.com.cn)。其他前沿模型也能接受多镜头提示词,但保真度各不相同——在投入一整份分镜之前,务必先在首次测试生成中验证切点。
时间戳在没有时间戳 API 的模型上有效吗? 有效——作为结构,而非承诺。在不解析时间戳的模型上写 “0:00–0:04” 依然有帮助,因为它强制了明确的时长、无缺口的节拍,以及一个模型必须纳入考虑的镜头边界。即使模型无法直接寻址时间线,你也在组织它的注意力。
为什么我的镜头在第一个之后还在一直动? 因为模型默认把摄影机运动当作片段的全局属性。修复方法是在每个镜头边界重置摄影机:每个镜头以明确的摄影机状态开头(”new setup — static”),说出运动的轴向和结束状态,永远不要指望上一镜头的指令会延续过来。
如何在多个镜头之间保持同一个角色? 三层:在主体出现的每个镜头中逐字重复的固定角色块;一条 “same in all shots” 的不变量条款;以及——在支持的地方——堆叠在提示词下方的参考图。完整的锚点格式和漂移检查清单见我们的角色一致性指南。
多镜头提示词比链接多次生成更好吗? 当片段较短、镜头共享同一地点与服装、且你想让模型自己处理切点时,它更好。当总长超过约 30 秒、某个镜头需要重拍、时长必须精确到帧,或场景差异大到一个连贯性场无法覆盖时,链式生成更好。请参见上面的决策表。
在单条多镜头提示词中应该避免哪些转场? 声音先行的转场(J 剪、L 剪)和装饰性转场(圆形划像、翻页)——在照片级真实风格上,模型经常无视它们。默认坚持硬切;想要融合时用叠化;需要命中时,用空间化描述的匹配剪辑和甩摇。
结论:写剪辑,而不是写场景
从”每条提示词一个片段”到”多镜头生成”的转变,改变了提示词的*本质*。它不再是场景的描述,而成为一份剪辑文档:镜头编号、时长分配、切点命名、每个边界处的摄影机运动重置,以及被声明为不变量的连贯性锚点——而不是可被重新随机的细节。
先从你的分镜原生形式开始——老式的”SHOT 1: WIDE”文档完全够用——然后跑完三个动作:给镜头编号并定时,把术语转换成方向性句子,把一切铺到逐秒的时间线上。在每个块之间加入明确的切点,逐字重复角色块和环境块,每次镜头变化时重置摄影机。上面有十二个模板可供复制。
然后,拆分要有意为之,而不是出于习惯:故事还装得进模型的连贯性场时,就用单次生成;一旦装不下,就立刻改用链式生成;介于两者之间的一切,用时间戳编辑。
从这里继续:
- Seedance 2.5 提示词:时间戳控制与多镜头结构 — 关于可按时间线寻址的提示词的深入解读。
- Kling 3.0 角色一致性 — 构建锁定身份的 6 镜头原生提示词。
- 产品转场效果视频提示词 — 把转场作为独立技术。
- Seedance 上的 Ref2V:用参考图保持分镜连贯 — 在多镜头提示词下堆叠参考帧。
- 电影级运镜 AI 视频提示词 — 逐镜头机位词汇目录。
- 2026 年自然的 AI 视频提示词公式 — 每个多镜头提示词继承的基础语法。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章