Seedance 2.5 多参考图像提示词:最大化利用 50 素材栈
速览
- Seedance 2.5 单次生成最多接受 50 个参考素材:30 张图片、10 个视频和 10 条音频——这是当前文生视频模型中最大的参考素材配额。
- 参考素材分为 7 种功能角色:角色(Character)、产品(Product)、环境(Environment)、风格(Style)、动作(Motion)、镜头(Camera)和音频(Audio)。
@tag语法(如@Image1、@Video3、@Audio2)允许你在提示词正文中为每个槽位贴标签并分配角色。- 8–12 个标注清晰的参考素材优于 30 个杂乱的上传文件——模型注意力层中槽位 1–5 承载最高的语义权重。
- 冲突规避即显式角色锁定:当两个参考素材可能影响同一要素(如光照与角色)时,明确指定胜出者。
- 以下 12 个即用模板涵盖品牌广告、角色一致性叙事、产品展示和音乐视频。
- 已在 Seedance 2.5 网页控制台和公开 API 上测试至 2026 年 10 月。
为什么 50 参考素材改变了一切
文生视频模型最初面世时,提示词就是全部产品。你写下”一位身穿红裙的女子在日落时分走过东京,电影感,24 帧”,然后寄希望模型已充分内化了关于东京、红裙和电影感光照的知识,从而渲染出接近的结果。效果如同老虎机碰运气。
Seedance 2.5 颠覆了这一工作流。你不再要求模型仅凭文字*想象*面孔、产品、地点或配乐,而是附加实际参考素材,并告诉模型哪个槽位掌管输出的哪个方面。单次生成最多可上传 30 张图片、10 个视频和 10 条音频——总共 50 个素材,因此得名”50 素材栈”。
三类素材各有不同用途:
| 素材类型 | 单次最大数量 | 典型用途 |
|---|---|---|
| 图片 | 30 | 角色面孔、产品照、风格定帧、环境底版、调色参考 |
| 视频 | 10 | 镜头轨迹、运动样本、对口型参考、预演片段 |
| 音频 | 10 | 环境底噪、画外音、音乐分轨、音效提示 |
这一转变是根本性的。提示词工程不再是”描述场景”,而是“路由参考素材”。你不再是一名文案,指望一串形容词能召唤出一张面孔;你是一名导演,为 50 人的剧组分配角色。
我们在 Seedance 2.5 网页控制台和公开 API 上的测试表明,提示词长度在 380–420 个标记的散文处趋于平稳——超过此量,模型开始忽略文本,转而以参考素材为准。模型将你的文本视为*路由表*而非剧本。参考素材承载视觉重量;文本决定当它们重叠时谁胜出。seedance2ai.net 的 50 参考素材指南 以不同的案例分析阐述了同样的洞察。
参考素材组:心智模型
50 个槽位并非 50 个同等的声音。它们归入七个功能角色,以角色来思考是对多参考素材工作流最大的升级。
| 角色 | 控制内容 | 最佳素材类型 | 槽位优先级 |
|---|---|---|---|
| 角色 | 各镜头中面孔、体型、服装的一致性 | 图片(面部特写、全身转向图) | 1 |
| 产品 | 物品识别、包装、品牌标识 | 图片(主图、多角度) | 2 |
| 环境 | 地点、背景底版、天气、时段 | 图片或视频底版 | 3 |
| 风格 | 调色、渲染风格、绘画感 vs. 摄影感 | 图片(截帧或情绪板) | 4 |
| 动作 | 主体在各帧间如何运动(步行循环、手势、舞步) | 视频片段 | 5 |
| 镜头 | 推拉、摇臂、斯坦尼康感、焦距 | 视频片段 | 6 |
| 音频 | 环境底噪、画外音、音乐、音效时机 | 音频文件 | 7 |
“按组思考,而非按槽位”这一规则意味着你不必上传 50 个独立文件;你为每个角色上传一到两个素材并贴标签,使模型知道哪个槽位承担哪个角色。我们测试中一个典型的品牌视频任务使用:
- 角色:2 张面部参考(正面 + 3⁄4 侧角)——槽位 @Image1、@Image2
- 产品:4 张主图——槽位 @Image3–@Image6
- 环境:2 张地点底版——槽位 @Image7、@Image8
- 风格:1 张氛围定帧——槽位 @Image9
- 动作:1 个运动样本——槽位 @Video1
- 镜头:1 个预演片段——槽位 @Video2
- 音频:1 条音乐底床 + 1 条环境音——槽位 @Audio1、@Audio2
总共 12 个参考素材——远低于 50 上限,且整齐归入七个角色。详细方案记录在 forvideo.ai 的 Seedance 2.5 参考素材指南中,该指南以一条 30 秒广告为案例,完整演示了”4 产品 + 2 地点 + 1 动作 + 1 音频”的分镜拆解。
@tag 语法
你上传的每个参考素材都有一个槽位编号——@Image1 到 @Image30、@Video1 到 @Video10、@Audio1 到 @Audio10。这些标签出现在提示词正文中你希望该参考素材*生效*的任意位置。模型的注意力偏向最早的槽位,因此槽位排序即优先级。
一个典型的标注提示词如下:
@Image1 — main character face (front view, neutral expression)
@Image2 — same character, 3/4 angle
@Image3 — product hero shot, white background
@Image4 — product in lifestyle context
@Image5 — modern office environment, golden hour
@Image6 — color palette reference, warm teal + amber
@Video1 — camera move: slow dolly-in 0–100cm, eye-level
@Audio1 — ambient bed, soft office hum
@Audio2 — music stem, upbeat corporate
PROMPT:
@Image1 and @Image2 are the same character, Dr. Maya Chen, mid-30s, wearing a navy blazer.
She stands in @Image5 (modern office, golden hour) and presents @Image3 (product on a pedestal).
Camera follows @Video1 (slow dolly-in) the entire scene.
Color grading pulled from @Image6 (warm teal + amber).
Audio bed: @Audio1 underneath the entire clip, @Audio2 fades in at second 8.
三点需要注意:
- 首块是标签映射表,不是散文。 它在模型”看到”参考素材之前声明每个参考素材*是什么*。
- 提示词正文内联使用标签,将每个参考素材导向其角色。
- 槽位 1(@Image1)留给身份最关键的素材——通常是主角的面孔。模型对其关注最强。
@tag 约定与 Seedance 2.5 官方提示词架构保持一致,详见 50 参考素材指南。部分第三方界面可能使用 #Image1 或 [img1];底层 API 为 @Image1。
质量优于数量
一个常见错误是贪多求全——一次性上传 30 张图片、10 个视频、10 条音频,”以防万一”。在我们的测试中,这是你能做的最糟糕的事。模型的注意力并非无限;它有一个预算,每增加一个参考素材就会被稀释。
8–12 个标注清晰的参考素材优于 30 个杂乱的素材。 原因如下:
- 每新增一个参考素材都会增加噪声。一张低质量的产品照放到 @Image12 不仅帮不上忙——它还会与 @Image3 形成竞争。
- 模型假设你上传的*所有*参考素材都相关。如果你同时上传厨房和海滩,可能得到一个”厨房-海滩”混合体。
- 槽位优先级会衰减。模型对槽位 1–5 的关注最强;据我们的非正式观察,槽位 20–30 的参考素材权重远低于槽位 1(Seedance 团队截至 2026 年 10 月尚未公布精确的注意力分布)。
经验法则:将槽位 1–5 优先分配给身份关键项(角色、产品、环境)。槽位 6–15 留给次要风格和动作。槽位 16–50 用于氛围、调色和音效——你不需要太精确的部分。
模型的能力上限(30 秒输出、4K 分辨率)记录于 mindstudio.ai,而 vidmuse.ai 的 Seedance 2.5 指南 中的平台对比明确指出,不同界面暴露的*参考素材配额*不同——部分暴露全部 50 个,另一些则限制为 20 个。
冲突规避
我们测试中大多数失败的生成并非由劣质参考素材导致——而是相互冲突的参考素材导致。两个参考素材都声称掌控同一要素,却没有关于谁胜出的指令。
常见冲突场景:
| 冲突 | 示例 | 解决方案 |
|---|---|---|
| 光照 | @Image1(角色)在柔和日光下拍摄;@Image5(环境)则是阴郁夜景底版 | 添加:”@Image1 lighting matches @Image5 palette” |
| 服装 | 角色参考为红衬衫;提示词却写”蓝色西装外套” | 重新拍摄或重新描述角色参考 |
| 镜头高度 | @Video1 是低角度镜头;@Video2 是平视预演 | 只使用一个镜头参考,或明确指定 “blend: low-angle start, eye-level finish” |
| 音频氛围 | @Audio1 节奏明快;@Audio2 情绪低沉 | 显式交叉淡入淡出:”@Audio1 plays 0–8s, @Audio2 takes over 8–30s” |
| 风格 | @Image6(风格参考)为绘画感;@Image5(环境参考)为摄影感 | 提升一个:”@Image6 style overrides all other grading” |
通用规则:当两个参考素材可能影响同一要素时,明确指定胜出者。 在提示词正文中加一行如 ”@Image5 lighting is authoritative; @Image1 inherits from it” 即可解决我们测试中的大多数冲突。
如需深入了解模型在架构层面如何解决这些冲突,请参阅 Picassoia 对 50 参考素材合成流程的分析。
12 个提示词模板
所有模板遵循相同的约定:顶部为标签映射表,下方为带标签的提示词正文。将你自己的参考素材替换到 @-槽位中即可。每一个都是经过测试的配置,而非纯理论。
品牌视频广告(3)
模板 1 —— 产品发布主推广告(15 秒)
@Image1 — product hero shot, white seamless background, front 3/4 angle
@Image2 — product detail (texture/material close-up)
@Image3 — lifestyle context: someone holding the product in a real environment
@Image4 — brand color palette reference
@Video1 — camera: slow 360° orbit around product, eye-level
@Audio1 — music: confident, minimal synth bed, builds over 15s
PROMPT:
The product in @Image1 rotates slowly on a white pedestal, lit to match @Image4 (brand palette).
Camera follows @Video1 (360° orbit) — never breaks the orbit.
At second 6, cut to a hand picking up the product from @Image3 angle.
Texture detail from @Image2 flashes for 1 second at second 10.
@Audio1 plays under the entire spot, peak volume at second 14.
模板 2 —— 创始人故事 / “关于我们”(30 秒)
@Image1 — founder face, front view, natural light
@Image2 — founder face, 3/4 angle, candid
@Image3 — workspace environment, day-in-the-life
@Image4 — product line-up shot
@Video1 — handheld previz clip, intimate, slight handheld wobble
@Audio1 — ambient bed: café + keyboard typing
@Audio2 — music: warm piano, swells at second 12
PROMPT:
@Image1 and @Image2 are the same founder, mid-40s, natural expression.
Open in @Image3 (workspace), founder looks up and greets camera.
Camera is handheld per @Video1 — never stabilizes to tripod.
Founder gestures toward @Image4 (product line-up) on a shelf at second 14.
@Audio1 runs 0–8s, @Audio2 crossfades in at second 8 and carries the rest.
模板 3 —— 季节性 / 广告变体(20 秒)
@Image1 — character from previous campaign, winter wardrobe
@Image2 — new seasonal environment (snowy street, holiday lights)
@Image3 — brand style frame from Q4 mood board
@Video1 — previz: tracking shot, walking forward, camera slightly ahead
@Audio1 — music: holiday-flavored version of brand theme
PROMPT:
Same character as @Image1, now in @Image2 (snowy holiday environment).
Color grading pulled from @Image3 — saturated reds and golds.
Camera moves with character per @Video1 (tracking, walking forward).
@Audio1 plays throughout, slight reverb tail in the last 4 seconds.
角色一致性叙事(3)
模板 4 —— 多镜头对话场景
@Image1 — Character A face, front view
@Image2 — Character A face, profile view
@Image3 — Character B face, front view
@Image4 — Character B face, profile view
@Image5 — café interior, two-person booth
@Video1 — previz: shot/reverse-shot pattern, medium close-ups
@Audio1 — ambient café bed
@Audio2 — music: light jazz undertone
PROMPT:
@Image1/@Image2 = Character A. @Image3/@Image4 = Character B.
Scene: @Image5 (café booth). A speaks first, B replies.
Camera uses shot/reverse-shot from @Video1.
@Audio1 throughout, @Audio2 ducks under dialogue at second 4 and 11.
模板 5 —— 角色跨越多个环境
@Image1 — character, front view, neutral
@Image2 — character, 3/4 angle
@Image3 — environment A: city street
@Image4 — environment B: rooftop
@Image5 — environment C: subway platform
@Video1 — motion: walking gait, casual pace
@Audio1 — music: continuous beat under montage
PROMPT:
Same character (@Image1/@Image2) appears in @Image3, @Image4, @Image5 sequentially.
Each environment: 5-second shot, walking per @Video1 motion.
@Audio1 threads through all three shots for continuity.
Wardrobe stays identical across all three — no costume changes.
模板 6 —— 衰老 / 延时角色弧线
@Image1 — character at age 20
@Image2 — character at age 40
@Image3 — character at age 60
@Image4 — environment: same location, three different decades of decor
@Video1 — slow zoom-out motion
@Audio1 — music: gentle, nostalgic, three movements
PROMPT:
@Image1 → @Image2 → @Image3 in sequence, same character aging.
@Image4 (environment) shifts subtly behind them to match the decade.
Camera slowly pulls back per @Video1 across the full 25 seconds.
@Audio1 has three movements that align with each age transition.
产品展示 / 电商(3)
模板 7 —— 单品 360° 演示
@Image1 — product, front
@Image2 — product, back
@Image3 — product, top
@Image4 — product, bottom
@Image5 — product in hand (scale reference)
@Video1 — camera: smooth 360° turntable
@Audio1 — light, modern music loop
PROMPT:
@Image1–@Image4 are the same product from four angles.
@Image5 establishes scale — a hand holds the product at second 1.
Camera orbits per @Video1, cycling through all four angles.
@Audio1 loops under the demo, no VO needed.
模板 8 —— 多品产品陈列
@Image1–@Image6 — six products from the same line, white background
@Image7 — brand palette reference
@Image8 — lifestyle context (someone using one of the products)
@Video1 — camera: dolly along a shelf of products
@Audio1 — music: upbeat, retail-friendly
PROMPT:
@Image1 through @Image6 displayed on a clean shelf, evenly spaced.
Color grading from @Image7 (brand palette).
Camera dollies past the shelf per @Video1, settling on @Image3 at second 12.
At second 15, cut to @Image8 (lifestyle use) for 5 seconds.
@Audio1 throughout.
模板 9 —— 产品对比(A vs B)
@Image1 — Product A hero shot
@Image2 — Product B hero shot
@Image3 — environment: neutral countertop
@Video1 — camera: side-by-side split-screen feel
@Audio1 — music: neutral, decision-aid tone
PROMPT:
@Image1 on the left, @Image2 on the right, both on @Image3 (countertop).
Camera holds a wide shot per @Video1, slight push-in at second 8.
At second 10, push in to @Image1 detail; at second 18, push in to @Image2 detail.
@Audio1 throughout, fades to silence at second 25.
音乐视频 / 对口型(3)
模板 10 —— 单人歌手对口型表演
@Image1 — artist face, front
@Image2 — artist face, 3/4 angle
@Image3 — environment: dimly lit studio, single key light
@Video1 — previz: artist performance clip, head-and-shoulders, microphone in frame
@Audio1 — full song (the track being lip-synced to)
PROMPT:
Artist from @Image1/@Image2 performs in @Image3 (studio, key light).
Lip-sync locked to @Audio1 — every syllable must hit.
Camera stays close per @Video1 (head and shoulders), occasional push-in.
No other characters appear. 30 seconds total.
模板 11 —— 多表演者 / 群戏
@Image1 — performer 1
@Image2 — performer 2
@Image3 — performer 3
@Image4 — stage environment, dramatic lighting
@Video1 — previz: choreographed group performance clip
@Audio1 — full song
PROMPT:
Three performers (@Image1, @Image2, @Image3) on @Image4 stage.
Choreography follows @Video1 (group performance reference).
Camera cuts between wide shot and individual close-ups every 4 seconds.
Lip-sync anchored to @Audio1 throughout.
模板 12 —— 电影叙事音乐视频
@Image1 — main artist, hero look
@Image2 — love interest / narrative partner
@Image3 — narrative environment 1 (urban rooftop)
@Image4 — narrative environment 2 (rainy street)
@Video1 — camera: cinematic dolly + crane combo
@Audio1 — full song with quiet intro and loud chorus
PROMPT:
Story unfolds across @Image3 and @Image4.
@Image1 and @Image2 are the only characters.
Camera moves cinematically per @Video1 — dolly in for verses, crane up for choruses.
@Audio1 drives the pacing: verses in @Image3, choruses in @Image4.
No lip-sync required — performance is narrative, not concert-style.
参考素材陷阱
以下是我们最常遇到的失败模式。它们均非灾难性,且均可避免。
质量问题。 一张模糊的 1024 像素产品照上传到 @Image3 会毒化后续所有输出。模型会放大参考素材,但不会修复它们。在我们的测试中,长边低于约 1500 像素的素材会明显降低输出质量。
冲突参考素材。 上文已述——两个参考素材声称掌控同一要素,却没有关于谁胜出的指令。务必在提示词正文中明确指定胜出者。
动作视频参考素材借用了主体而不仅是镜头运动。 一个常见错误:将一段音乐视频片段上传为 @Video1,却期望模型只使用其*镜头运动*。模型还会借用该片段中的表演者、服装和光照。如果你只想要镜头轨迹,请使用预演或动画分镜片段——而非成品表演片段。
风格参考素材压制一切。 一张强烈的 @Image6(风格定帧)可能覆盖你精心挑选的角色参考。如果风格与角色冲突,将风格降级到槽位 12 以上或完全移除。
音频参考素材缺少时间标记。 将一首 3 分钟的歌曲上传到 @Audio1 并期望模型自行决定使用其中哪 30 秒是行不通的。请将音频文件预剪到与输出一致的长度,或在提示词中指定时间范围:”@Audio1 plays 0:30–1:00 of the uploaded track.”
过度依赖上限 50。 即便填满全部 50 个槽位,模型的有效注意力预算上限约为 12–15 个有意义的参考素材。超过槽位 15 的部分,统计上就是噪声。
单参考素材的惯性思维。 从单图生视频工作流迁移过来的团队常常只上传一张”好”图,期望它包打天下。在 50 参考素材的世界中,这等于让 49 个槽位闲置。将你的视觉简报拆解——面孔、产品、环境、风格、动作、镜头、音频——并将每项分配到各自的槽位。
常见问题
我到底该用多少个参考素材? 在我们的测试中,8–12 个标注清晰的参考素材始终优于 20–30 个上传项。身份关键素材(角色、产品、环境)应放在槽位 1–5。其余槽位用于次要风格和氛围。
@Image1 总是代表主角吗? 它代表*你决定它代表的含义*——槽位 1 承载最多注意力,因此应分配给身份最关键的素材。对于纯产品广告,槽位 1 可能是产品主图而非面孔。
同一参考素材可以在多个提示词中重复使用吗? 可以。槽位编号仅在单次生成内有效;每次新提示词需重新上传或重新关联参考素材。部分 API 客户端支持参考素材库——上传一次,反复使用。
如果上传了参考素材却没有在提示词中标记它,会怎样? 模型仍会尝试使用它,但置信度很低——它会作为背景噪声或部分影响出现。请始终在提示词正文中为每个上传的参考素材贴标签。
Seedance 2.5 支持负面参考素材(指定避免项)吗? 不支持专门的语法。变通方法是上传一个参考素材并描述需要*避免的特征:”@Image7 is what the background must NOT look like — flat, gray, sterile.“* 此方法效果不稳定。
这与单参考素材的图生视频工作流有何不同? 单参考 I2V 将一张图片视为完整的视觉简报。Seedance 2.5 的多参考素材栈允许你将视觉简报*分区*到多张图片、视频片段和音频文件中——更接近电影剧组交接,而非单一静帧。
50 参考素材上限在所有平台上都一样吗? 模型规格一致(30 图、10 视频、10 音频),但界面实现各异。官方网页控制台暴露全部 50 个;部分第三方工具限制为 20 个或仅暴露图片参考。vidmuse.ai 的 Seedance 2.5 指南 提供逐平台对比。
结论
50 参考素材栈是 Seedance 2.5 最重要的功能。它将创作瓶颈从*描述*场景转移到*路由*素材——而路由是大多数团队比描述更擅长的事。八到十二个标注清晰的参考素材,按七个角色(角色、产品、环境、风格、动作、镜头、音频)组织,几乎能胜任任何短视频任务。
如需了解 Seedance 2.5 除多参考素材以外的更广泛提示词策略,请参阅我们的 Seedance 2.5 模型指南 以及关于 多模态参考视频提示词 的姊妹篇。如果你专注于参考到视频流水线,Ref2V 配套文章 更深入地探讨 API 端工作流。针对角色一致性场景,图生视频角色锁定提示词 以及我们的 2026 年 AI 角色一致性提示词 进一步完善了全貌。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章