多模态参考视频提示词:将图像、视频和音频作为输入
速览
- 多模态参考视频提示词将文本、图像、视频和音频输入统一编排,使每个输入各司其职地掌管最终画面的某一方面(主体、动作、氛围、节拍)。
- 单模态提示词在需要同时锁定角色 + 镜头轨迹 + 环境音 + 视觉风格时就会触顶;只有多输入编排才能同时稳住这四个维度。
- 三种主要输入模式为 图像(主体、风格、调色)、视频(动作、编排、镜头运动)和音频(节奏、氛围、人声、音效),每种都必须在提示词中分配明确角色。
- 最难的部分不是上传——而是路由:告诉模型哪个输入掌管主体,哪个掌管动作,哪个掌管光照,让参考素材之间互不抢戏。
- 2026 年模型支持矩阵:Seedance 2.5 支持多达 50 个素材参考,Veo 3.1 接受图像 + 音频,Kling 3.0 使用元素槽位,Wan 3.0 支持起止帧加音频。
- 跨模态锚定(为每个参考素材显式分配角色)是让角色一致性 + 镜头编排协同工作的关键;缺少它,模型只会挑一个参考素材而忽略其余。
- 使用针对多模态的负面提示词基线(如”图像与视频参考之间不出现风格串扰”、”物体不得跨镜头漂移”),可抑制常见的失败模式。
2026 年视频生成中”多模态”的真正含义
多模态参考视频提示词指模型接收多于一种输入信号的提示词——通常包括一段文本提示、一张或多张图像、一段或多段参考视频,以及可选的一段音频文件——并且提示词中明确声明每个输入负责什么。
听上去简单。实际上,这是自 2024 年图像条件引入以来视频生成领域最大的可用性跃迁。在 2025 年末之前,大多数生产级视频工作流要么是纯文本,要么是文本加单张图像。这层天花板随处可见:你可以从一张图锁定人脸,但无法同时锁定一段参考片段的镜头运动;你可以描述咖啡馆的环境音,但无法将其锚定到一段真实的音频文件;你无法在 12 个镜头的序列中保持角色造型一致,因为每次重新生成都会重摇服装。
2026 年的模型栈——Seedance 2.5、Veo 3.1、Kling 3.0、Wan 3.0——打破了这一天花板,能够并行接收多个参考素材并让每个素材承载独立信号。但哪个输入代表什么并不会由模型替你决定。如果你喂给它一张女性肖像和一张电影剧照,再写”匹配这种美学”,它会在编码空间里挑一个读起来更强的参考素材,忽略另一个。提示词的工作就是消除这种歧义。
三股趋势汇合,使多模态从例外变为默认。第一,参考图像支持现已成为各家模型的标配——即便是低端玩家也至少接受一张图像。第二,参考视频槽位已从实验性走向生产级:Seedance 2.5 的 ref2v 通路和 Veo 3.1 的参考视频槽位均作为稳定 API 上线,不再是研究预览。第三,音频条件已跨越门槛,从新鲜事物变为可围绕其搭建真实场景的能力,尤其是 Veo 3.1 的环境与人声通道以及 Wan 3.0 的节奏锚定槽位。
这一切的代价是提示词复杂度的上升。纯文本提示词可以只是一句话。多模态提示词更接近一份附带元数据的分镜表:哪个输入是主体,哪个是光照,哪个是镜头,哪个是节拍。本文后半部分正是围绕这一核心展开。
这就是多模态参考提示词的完整学科。后续章节将介绍输入模式、路由难题、按模型划分的支持矩阵、锚定技术以及 12 个可即用的模板。
如需了解单输入模式的背景知识,请参阅我们的 Seedance 2.5 提示词指南 以及关于 Seedance 多参考图像提示词 的姊妹篇,后者覆盖了纯图像参考下的同一路由学科。
三种输入模式
每一条多模态视频提示词都是三种输入模式(外加文本)的组合。理解每种模式的原生信号是正确路由的前提。
1. 图像输入——主体、构图与风格
图像参考为生成贡献三样东西:主体(面孔、体型、造型、道具)、构图(取景、景深、前景/背景层次)以及视觉风格(调色、光照方向、质感、年代感)。支持图像参考的模型会将其编码进视觉-语言嵌入空间;接着由文本提示词引导该嵌入应主导上述三个面向中的哪一面。
一个实用写法:
Image #1 (portrait.jpg): woman, late 20s, red trench coat, soft window light.
Text: "Use Image #1 as the subject reference only. Do not copy its lighting. Place her in a rainy neon alley, mid-stride, three-quarter framing."
此处图像被路由为”仅做主体”,文本显式覆盖其光照。若无此覆盖,模型会复制图像的整体情绪氛围,你的文本提示词将沦为装饰。
图像参考也是跨多个镜头锁定角色一致性的最干净途径;关于该工作流的深入讨论,请参阅 2026 年 AI 角色一致性提示词 以及 图生视频角色锁定提示词。
2. 视频输入——动作、镜头轨迹、编排
视频参考贡献的是任何图像都无法提供的时间信息:镜头运动、角色动作、场景节奏、镜头行为以及剪辑节奏。当你上传一段参考片段时,模型会编码其运动矢量并将其用作先验。
Video #1 (whip_pan.mp4): handheld dolly shot panning across a crowded market.
Text: "Match the camera movement and rhythm of Video #1. Subject is a chef in white apron (described in text). Do not copy the market environment — render a quiet kitchen instead."
这里的陷阱是过度复制。一段视频参考在其像素中同时承载环境、主体、光照和动作。如果你的文本提示词不指明应保留哪些面向,模型通常会默认保留动作而丢弃环境——但在某些模型中,环境会占上风。永远要消除歧义。
针对专精”视频作为输入”的模型(Seedance 2.5 的 ref2v 以及 Veo 3.1 的参考视频槽位),这是整套工具箱中最强大的单一模式。Ref2V 配套文章 详细讲解了其机制。
3. 音频输入——节奏、氛围、人声、节拍同步
音频输入是三者中最新的一种,支持度也最不一致。截至 2026 年 10 月,仅 Veo 3.1 和 Wan 3.0 原生接受音频文件(Seedance 2.5 在 2.5 更新中新增了一个实验性音频条件槽位,Kling 3.0 仅将其用于时序而非内容)。音频贡献四类信号:环境音底床(咖啡馆、风声、室内底噪)、节奏/节拍(模型可据此同步动作)、对白/人声(部分模型可对口型)、音乐提示(可驱动转场)。
Audio #1 (rain_loop.wav): steady rainfall on a tin roof, no music.
Text: "Generate 8 seconds of a man walking through a forest at night. Match the rhythm of Audio #1's rainfall cadence — slow footsteps in sync with each drip cluster. No dialogue."
这是多模态协同工作最干净的范例:单条音频文件完成的节奏锚定,用文字根本无法描述。
关于 Veo 3.1 的完整音频栈,请参阅 Veo 3.1 原生音频 4K。
模型支持矩阵(2026 年 10 月)
并非所有模型都接受所有模式。下表汇总了四大主流视频模型各自接受哪些信号,以及每个输入控制什么。
| 模型 | 图像参考 | 视频参考 | 音频输入 | 参考素材上限 | 路由原语 |
|---|---|---|---|---|---|
| Seedance 2.5 | 是 | 是 | 实验性(仅节拍) | 跨所有模式最多 50 个素材参考 | 按素材的角色标签([subject]、[motion]、[style]) |
| Veo 3.1 | 是 | 是 | 是(环境 + 人声 + 节拍) | 通常为 3 图 + 2 视频 + 1 音频 | 参考槽位顺序 |
| Kling 3.0 | 是(元素槽位) | 无原生视频参考 | 否 | 最多 4 个元素槽位 | 元素槽位分配 |
| Wan 3.0 | 是(起帧 + 末帧) | 隐式(通过关键帧插值) | 是(环境 + 节拍) | 2 帧 + 1 音频 | 起止关键帧 + 音频条件 |
上表的资料来源:Seedance 2.5 文档综合自 mindstudio.ai 的 Seedance 2.5 功能详解 以及 seedance2ai.net 的 50 参考素材指南;Veo 3.1 能力引自 DeepMind 官方 Veo 模型页面 和 Veo API 手册。Kling 3.0 和 Wan 3.0 两行反映的是截至 2026 年 10 月在公开工具中的观察行为。
这张矩阵有两点特别突出。第一,Seedance 2.5 是组内唯一将参考素材数量扩展到个位数以上的模型——50 素材上限非同寻常,解锁了复杂的多元素场景。第二,Veo 3.1 是唯一拥有干净、生产就绪音频槽位的模型,它接受真实声音文件而非仅仅是节奏元数据。
路由难题
多模态参考视频提示词中最难的部分不是上传步骤。而是路由:决定哪个输入掌管输出的哪个方面,并在提示词中写明这一决定。
天真的失败模式看起来像这样:你上传一张角色肖像,同时上传一张《银翼杀手》的剧照。你写”赛博朋克女子穿行霓虹市场”。模型确实生成了结果,但搞不清楚女子的面孔是来自你的肖像,还是光照是来自你的参考剧照。你重新生成四次,得到四个不同的答案。
实际发生的事情是:模型把两个参考素材都当作等权重的风格先验。在没有显式角色分配的情况下,它每次生成时会随机采样其中一个。
解决办法是为每个参考素材在提示词中贴一个角色标签。具体语法因模型而异:
- Seedance 2.5:使用方括号标签,如
[subject=portrait.jpg]、[lighting=blade_runner_still.jpg]、[motion=walking_loop.mp4]。模型为每个槽位独立加权。 - Veo 3.1:参考素材上传到有序槽位中。第一个图像槽位视为主主体;后续槽位按顺序加权。文本提示词可显式降级某槽位(”Image #2 is style-only — ignore its composition”)。
- Kling 3.0:每个元素槽位都是一个有标签的实体。上传时给槽位贴标签(”character”、”outfit”、”prop”)。
- Wan 3.0:起帧和末帧默认角色无歧义;音频槽位同样无歧义。
如果在不强制分配角色的模型(Kling、Wan)上跳过角色分配,你会得到随机采样。而在强制分配的模型(Seedance、Veo)上,未分配的槽位默认为”style”——可能符合也可能不符合你的本意。
一个可靠的启发式:在描述场景之前先在提示词中写入角色分配。以”Use Image #1 as subject reference, Image #2 as lighting reference, Video #1 as camera motion”开头的提示词,在任何支持这些输入的模型上都能正确路由,无论槽位顺序如何。
还有一种次级失败模式值得指明。即使角色分配正确,模型偶尔也会*跑偏*:一个被标记为主体专属的图像,在一次糟糕的重新生成中会将其光照泄漏进最终帧;或一段被标记为动作专属的视频会将其调色渗入输出。这不是提示词层面的路由失败——而是模型内部的编码器失败。缓解方法有二。第一,编写紧凑的负面提示词(见下文)。第二,对每个镜头运行数次重新生成并挑选最干净的输出。多模态工作流并不像偶尔能做到的纯文本提示词那样是确定性的;平均而言,每个可用片段预算 3 到 5 次重新生成。
跨模态锚定技术
跨模态锚定是让单个参考素材在生成中扮演多种角色、或让两个不同参考素材锁定同一镜头不同面向的技术集合。
技术一:视频作为动作模板,图像作为主体
最常见的锚定模式。你有一段展示目标镜头运动的参考视频,以及一张应出现在镜头中的角色的肖像。若无锚定,模型会用视频中的角色替换。有了锚定,你可以告诉模型:
Video #1 [motion only]: handheld dolly-in from wide to medium close-up over 4 seconds.
Image #1 [subject only]: woman in red trench coat.
Text: "Render Video #1's camera move. Place Image #1's character in the frame. Background: empty warehouse."
这正是 Seedance 2.5 所设计的模式——Ref2V 工作流 正是在此基础上构建的。
技术二:图像作为风格,视频作为节奏
反过来:你想要一张图像的*观感*(一幅卡拉瓦乔的画),但要一段参考视频的*节奏*(一段 Derek Cianfrance 风格的长慢镜头)。若不显式贴角色标签,模型会字面复制图像(一幅静止画作)并完全丢弃视频。
Image #1 [style only]: Caravaggio's "The Calling of Saint Matthew", use only its color palette and chiaroscuro lighting.
Video #1 [motion only]: 8-second slow push-in with no cuts.
Text: "A modern executive sits at a desk. Light her as Image #1. Move the camera as Video #1."
技术三:音频作为节奏锚,图像作为风格与主体
针对 Veo 3.1 和 Wan 3.0,你可以将场景节奏钉在音频文件的节拍上,同时用图像参考保留视觉风格和角色:
Image #1 [subject + style]: young man in 1970s suit, saturated Kodachrome palette.
Audio #1 [rhythm only]: drum loop at 92 BPM.
Text: "Image #1's character walks across a motel parking lot at night. Match his footstep cadence to Audio #1's kick drum. No dialogue."
技术四:堆叠元素锚定(仅 Seedance 2.5)
Seedance 2.5 的 50 参考素材槽位让你一次性为多个小型参考素材分配不同角色。Seedance 多参考图像提示词文章 深入介绍了这一点——简短版本是:你可以在同一次生成中分别为角色面孔、角色造型、道具、背景、光照方向和镜头运动各自钉上一个参考素材。
12 个提示词模板
以下模板开箱即用:复制提示词,填入引用的文件,运行即可。每一项均标注了主要测试所用的模型;大多数只需稍改语法即可跨模型使用。
文本 + 1 张图像(3 个模板)
模板 1——从肖像锁定角色
Model: Veo 3.1 (also runs on Seedance 2.5, Kling 3.0)
Input: 1 image (character_portrait.jpg)
Role: Image = subject only. Text owns everything else.
[Reference: character_portrait.jpg] Subject reference only.
A woman in her late 20s walks through a rain-soaked Tokyo alley at night.
She wears a red trench coat over a white shirt, as in the reference.
Camera: handheld, slightly low angle, follows her from the side.
Lighting: neon signage from above, puddles reflecting pink and teal.
Duration: 6 seconds. Aspect 16:9. No dialogue. No cuts.
模板 2——从画作迁移风格
Model: Wan 3.0 (also runs on Veo 3.1)
Input: 1 image (oil_painting.jpg)
Role: Image = color palette and texture only. Text owns composition and motion.
[Reference: oil_painting.jpg] Use only its color grading and brushstroke texture.
A chef in a white apron plates a dish in a quiet kitchen.
Adopt the reference's muted ochre-and-slate palette and visible canvas grain.
Camera: static medium shot, gentle rack focus from hands to face.
Duration: 5 seconds. Ambient sound only.
模板 3——主图衍生产品镜头
Model: Seedance 2.5
Input: 1 image (product_hero.jpg)
Role: Image = subject + composition. Text adds motion and lighting.
[Reference: product_hero.jpg] Subject and framing reference.
Render the bottle from the reference on a rotating turntable.
Slow 360-degree orbit over 8 seconds.
Background: deep navy gradient, single soft top light catching the glass.
No text overlays. No hands.
文本 + 1 段视频(3 个模板)
模板 4——复用镜头轨迹
Model: Seedance 2.5
Input: 1 video (dolly_clip.mp4)
Role: Video = camera motion only. Text replaces subject and environment.
[Reference: dolly_clip.mp4] Camera motion template only.
Match the exact dolly-in trajectory and speed of the reference.
Subject: a young boy running through a wheat field at golden hour.
Replace the reference's environment entirely.
Duration: 5 seconds. No dialogue.
模板 5——复用节奏与韵律
Model: Veo 3.1
Input: 1 video (long_take.mp4)
Role: Video = pacing only (no environment, no subject).
[Reference: long_take.mp4] Use only the pacing — single continuous shot, no cuts, slow acceleration over 6 seconds.
Subject: a chess match between two elderly men in a park.
Do not copy the reference's setting.
Lens: 50mm equivalent, eye level.
Sound: park ambience, distant traffic.
模板 6——复用舞蹈编排
Model: Seedance 2.5
Input: 1 video (dance_clip.mp4)
Role: Video = choreography only. Subject and wardrobe from text.
[Reference: dance_clip.mp4] Choreography reference only.
Copy the dance sequence beat-for-beat.
Two dancers in matching black outfits perform the routine on an empty soundstage.
Camera: wide static shot to capture full body.
Lighting: single overhead spotlight, hard shadows.
Duration: 8 seconds. No music in audio (we will score in post).
文本 + 图像 + 视频(3 个模板)
模板 7——主体来自图像,动作来自视频
Model: Seedance 2.5 (canonical use case)
Input: 1 image (character.jpg), 1 video (walk_loop.mp4)
Role: Image = subject. Video = motion. Text owns environment and lighting.
[Reference: character.jpg] Subject only — face, hair, wardrobe.
[Reference: walk_loop.mp4] Motion only — match the walk cycle and arm swing exactly.
Place the subject in a crowded train station at rush hour.
Camera: tracking shot from the front, slight low angle.
Lighting: fluorescent overhead, mixed with daylight from entrance.
Duration: 6 seconds. Ambient sound: station PA, footsteps, distant train.
模板 8——风格来自图像,节奏来自视频,主体来自文本
Model: Veo 3.1
Input: 1 image (film_still.jpg), 1 video (long_lens.mp4)
Role: Image = color grading and lens choice. Video = pacing. Text owns subject.
[Reference: film_still.jpg] Color grading and lens only — adopt its desaturated teal-and-orange palette and shallow depth of field.
[Reference: long_lens.mp4] Pacing only — single 8-second shot, no cuts, slow lateral push.
Subject: a journalist interviewing a source in a dimly lit bar.
Do not copy the reference video's setting or subjects.
模板 9——多图像参考(风格 + 主体)+ 视频(镜头)
Model: Seedance 2.5
Input: 2 images (face.jpg, wardrobe.jpg), 1 video (camera_move.mp4)
Role: face = subject identity. wardrobe = costume. video = camera move.
[Reference: face.jpg] Subject's face only.
[Reference: wardrobe.jpg] Wardrobe and accessories only — red dress, gold earrings, as shown.
[Reference: camera_move.mp4] Camera trajectory only — match the slow crane-up.
The subject walks through an art gallery, pausing at a painting.
Lighting: warm gallery spots, neutral walls.
Duration: 7 seconds.
文本 + 图像 + 视频 + 音频(3 个模板)
模板 10——完整多模态:角色、动作、环境音
Model: Veo 3.1 (canonical four-mode prompt)
Input: 1 image (char.jpg), 1 video (pace.mp4), 1 audio (rain.wav)
Role: image = subject, video = pacing, audio = ambient bed and rhythm anchor.
[Reference: char.jpg] Subject only — woman in her 50s, grey hair, beige coat.
[Reference: pace.mp4] Pacing only — match the slow lateral track over 6 seconds.
[Reference: rain.wav] Ambient sound and rhythm anchor — match her walking cadence to the rainfall's intensity curve.
She walks through a park at dusk during steady rain.
No dialogue. No music.
模板 11——锁定编排的音乐视频镜头
Model: Seedance 2.5 (with experimental audio)
Input: 1 image (performer.jpg), 1 video (choreo.mp4), 1 audio (track.wav)
Role: image = performer identity, video = choreography, audio = beat-sync and lyrics.
[Reference: performer.jpg] Subject only — keep the performer's face and signature outfit.
[Reference: choreo.mp4] Choreography only — copy the routine's structure and timing.
[Reference: track.wav] Beat-sync and lip-sync — match movement to kick drum on beats 1 and 3.
Lip-sync to the chorus's vocal melody.
Background: dark soundstage with three rotating spotlights.
Duration: 12 seconds. Single shot.
模板 12——锁定视觉与环境音的纪录片访谈
Model: Veo 3.1
Input: 1 image (subject.jpg), 1 video (interview_style.mp4), 1 audio (room_tone.wav)
Role: image = subject, video = interview framing and gesture pacing, audio = room tone.
[Reference: subject.jpg] Subject identity only — older man with glasses, navy sweater.
[Reference: interview_style.mp4] Framing and gesture pacing only — medium close-up, subject nods on the third beat of each phrase.
[Reference: room_tone.wav] Room tone only — quiet indoor ambient with occasional HVAC hum.
The subject sits in a leather chair, speaking directly to camera about his career.
No background music.
Duration: 8 seconds.
多模态负面提示词基线
单模态视频生成有其自身的负面提示词词汇。多模态生成多出一层,因为参考素材之间会以特定的、反复出现的方式互相冲突。以下负面基线是我们用于每一条多模态提示词的标配:
Universal multimodal negatives:
- no style bleed between image and video references
- no subject shift across shots (character must remain identical to Image #1 in every frame)
- no lighting conflict between reference and text description
- no environment copy from video reference (unless explicitly assigned)
- no costume drift between reference image and generated frames
- no camera move that contradicts the video reference's trajectory
- no audio sync drift between motion and rhythm input
将以上内容加入每次多输入生成提示词的负面部分。它们应对最常见的失败模式——参考素材之间争抢同一面向的控制权。
按模型补充:
- Seedance 2.5:
no reference priority inversion(当某槽位泄漏进另一槽位的角色时)。 - Veo 3.1:
no off-slot reference contamination(当 Image #2 影响主体而非风格时)。 - Kling 3.0:
no element slot bleed(当道具槽位影响角色时)。 - Wan 3.0:
no start-end frame averaging(当模型模糊关键帧而非插值时)。
常见问题
什么是多模态参考视频提示词? 一种组合多种输入类型的提示词——通常是文本加上一张或多张图像、参考视频和/或音频文件——并明确说明每个输入在生成视频中所控制的内容。
2026 年哪些模型支持多模态参考输入? Seedance 2.5、Veo 3.1、Kling 3.0 和 Wan 3.0 都接受某种组合的图像、视频和音频参考。Seedance 2.5 拥有最大的参考素材预算(最多 50 个素材);Veo 3.1 拥有最完善的音频集成;Kling 3.0 使用元素槽位进行纯图像组合;Wan 3.0 使用起止关键帧加音频条件。
如何防止参考素材之间互相冲突? 在提示词中为每个参考素材分配明确的角色(主体、光照、动作、风格、节奏)。如果支持角色标签,使用它们。如果不支持,则在描述场景前用纯文本写出角色分配。见上文”路由难题”一节。
我能否同时用一段视频参考来提供动作和环境? 技术上可以,但不应如此。如果你两者都想要,请提供两个参考——一个标记为动作,动作,一个标记为环境,环境。模型对单角色参考的处理远比多角色参考可靠。
音频参考能用于人声和对白吗? 仅在 Veo 3.1 的人声槽位中可行,且仅限短句。截至 2026 年 10 月,对口型在大多数模型上仍不可靠。对白密集的场景,请先生成无声画面,再在后期配音。
一次能同时用多少个参考素材? 取决于模型。Seedance 2.5 上限为 50 个素材参考;Veo 3.1 通常接受 3 图 + 2 视频 + 1 音频;Kling 3.0 上限为 4 个元素槽位;Wan 3.0 为起止帧 + 音频。更多参考素材并不总是意味着更好的输出——大多数生产级提示词使用 2 到 4 个参考素材,其余交给文本。
最常见的多模态提示词错误是什么? 忘记分配角色。上传参考素材却不告诉模型每个控制哪个面向,会导致随机采样和重新生成不一致。务必先写角色分配。
结论
多模态参考视频提示词的本质是路由——告诉模型哪个输入掌管输出的哪个面向。一旦内化这一点,其余都是机械操作:选择你的参考素材,为每个分配角色,描述场景,加上合适的负面基线。
上面的 12 个模板只是起点。模型版图会持续变化——Seedance 2.5 的 50 素材槽位很可能在 2027 年被竞争对手追平,音频条件将变得普及,视频对视频的参考将取代部分图像参考。路由学科保持不变。
如需阅读邻近工作流的更多内容:
- Seedance 多参考图像提示词 — 大规模纯图像参考组合
- Seedance 2.5 提示词 — 模型指南与基础提示词模式
- Veo 3.1 原生音频 4K — Veo 音频栈详解
- Ref2V 配套文章 — 视频作为输入的深度探讨
- 2026 年 AI 角色一致性提示词 — 跨镜头保持角色稳定
- 图生视频角色锁定提示词 — 首帧角色锚定
本文引用的外部资料:
- Seedance 2.5 功能详解 — mindstudio.ai
- Seedance 2.5 50 参考素材指南 — seedance2ai.net
- Veo 模型页面 — DeepMind
- Veo 手册 — Google AI Studio
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章