Seedance 2.5 提示词库:2026 年完整指南
摘要
- Seedance 2.5 由 ByteDance 于 2026 年 7 月 31 日 发布,可根据文本、图像、视频和音频输入生成最长 30 秒的原生 4K 视频。
- 该模型通过
@标签原生支持最多 30 张图像、10 个视频片段和 10 个音频文件 作为多模态参考。 - 最可靠的提示公式为:
[主体] + [动作] + [场景] + [镜头] + [光线] + [风格] + [声音]。 - Seedance 2.5 擅长 因果逻辑链 —— 描述因果关系的提示词(例如”蜡烛闪烁是因为门被打开”)能产出最连贯的结果。
- 本提示库涵盖 30+ 即拷即用模板,分为电影叙事、产品展示、角色驱动和抽象/艺术四大类。
- 负向提示词 和 时间戳分段控制 是区分业余输出与专业级画面的两个进阶利器。
什么是 Seedance 2.5?
Seedance 2.5 是 ByteDance 的旗舰文本转视频系统,于 2026 年 7 月 31 日 通过官方 Seed 团队博客(seed.bytedance.com/zh/blogs/)正式发布,并由中国新闻网(chinanews.com.cn)深入报道。它接替了 Seedance 1.0 和更早的 1.5 Pro 版本,是首个可在单次推理中生成 30 秒原生 4K 片段 并广泛可用的生成器(new.cgvisual.com)。
三项规格使 Seedance 2.5 脱颖而出:
- 原生 4K 分辨率,30 秒时长。 此前的生成模型要么上限为 1080p,要么从较低分辨率的潜变量进行上采样。Seedance 2.5 原生生成 2160p 帧(chinanews.com)。
- 单条提示中最多 50 个多模态参考 —— 30 张图像、10 个视频片段和 10 个音频文件,可通过
@标签寻址(new.qq.com)。 - API 于 2025 年 8 月上线,支持生产集成(chinanews.com.cn)。
公开发布还解锁了 多镜头叙事生成 —— 模型可在单条提示中规划场景转场,将每个段落视为因果关联的街节,而非孤立的片段。
在 videosprompt.org 编辑实验室的测试中,我们一致发现 Seedance 2.5 在长篇叙事连贯性方面优于竞品,但前提是遵循特定的结构化公式。泛泛的”电影感城市无人机镜头”提示效果平平;而编码了因果关系和参考结构的提示则可达到专业水准。
Seedance 2.5 提示公式
经过对数百次生成的评估,推荐将以下七部分公式作为每条提示的基础:
[主体] + [动作] + [场景] + [镜头] + [光线] + [风格] + [声音]
| 组件 | 作用 | 示例 |
|---|---|---|
| 主体 | 焦点实体是谁或什么 | “一位年轻的大提琴手” |
| 动作 | 主体做了什么,最好包含因果关系 | “窗帘被掀起,因为她打开了窗户” |
| 场景 | 何地何时 | “黎明时分阳光明媚的巴黎阁楼内” |
| 镜头 | 拍摄构图、镜头、运动 | “缓慢推轨,35mm 变形宽银幕,浅景深” |
| 光线 | 光照方向、色温 | “金色时刻的暖色背光从窗户射入” |
| 风格 | 视觉处理或参考 | “电影感,Kodak Vision3 5219 胶片模拟” |
| 声音 | 音频描述或 @audio 引用 | “轻柔的翻页声,远处教堂钟声” |
使用该公式的完整提示:
A weathered sailor in a wool peacoat stands at the bow of a fishing trawler,
gripping the railing because the sea swells with an incoming storm.
The scene is the open Atlantic at first light, fog drifting across gunmetal waves.
Camera: medium close-up, slow handheld push-in, 50mm lens, slight roll.
Light: cold blue-grey overcast with a single warm rim light from the wheelhouse.
Style: cinematic realism, ARRI Alexa LF emulation, desaturated teal grade.
Sound: wind gusting across the mic, gull halfrated, creaking timber hull.
该公式并非僵化不变 —— 各部分可以合并(如”月光下的冷调电影音乐”)或省略 —— 但每个角色都促使你做出明确选择,而非依赖模型自行填补空白。模型往往默认采用中性光照和通用运动方式;而完整指定七部分可产出明显更优的画面。
因果逻辑链
Seedance 2.5 最被低估的能力是 因果逻辑:模型可以对物体、角色和环境之间的因果关系进行建模。这是 ByteDance 2026 年研究路线图的重点之一(seed.bytedance.com),也是 Seedance 与将视频视为不相关帧序列的竞品之间最显著的区别。
因果提示词先指明 结果,再用 because、which causes、resulting in、as a result of、triggers 等连接词附上 原因。
因果逻辑的三大原则
| 原则 | 描述 | 提示摘录 |
|---|---|---|
| 物体因果 | 一个物体物理上导致另一个物体发生变化 | “花瓶倾倒是因为猫从它身边走过” |
| 环境因果 | 环境对主体作出反应 | “雾气消散是因为太阳从山脊后升起” |
| 情感因果 | 主体表情因情境而变化 | “她微笑是因为女儿正穿过田野向她跑来” |
嵌入因果关系的示例提示:
A toddler in a red raincoat jumps into a puddle because she heard the first
thunderclap. The splash ripples outward across the asphalt, soaking her boots.
Camera: low-angle 24mm, slow-motion 120fps, tracking right.
Light: overcast afternoon, soft diffused fill from a grey sky above.
Style: photojournalistic, Fujifilm X-T5 film emulation, slight grain.
Sound: splash, distant thunder, child's laughter echoing off brick walls.
在我们的测试中,因果提示词能有效减少时间伪影 —— 模型会更为审慎地规划”前置”帧,因为它必须满足原因条件,从而迫使”后续”帧进入连贯状态。非因果提示词(如”孩子在水坑中踩水”)则会产生更多形变与身份漂移。
多模态参考标签(@image、@video、@audio)
Seedance 2.5 的 API 和网页界面在每条提示中接受最多 50 个参考素材:30 张图像、10 个视频片段 和 10 个音频文件(new.qq.com)。每个素材通过 @tag 语法寻址,将参考与提示文本中的特定元素绑定。
标签语法
| 标签 | 用途 | 使用次数限制 |
|---|---|---|
@image1、@image2… |
将上传的图像作为视觉来源参考 | 最多 30 |
@video1、@video2… |
参考上传的视频片段以获取运动/风格 | 最多 10 |
@audio1、@audio2… |
参考上传的音频文件作为声音层 | 最多 10 |
提示中的标签是位置式的 —— @image1 指的是上传的第一张图像,与文件名无关。模型将标记的素材绑定到提示中 最近的名称。
包含多模态参考的示例
A ceramic teacup @image1 sits on a slate countertop, steam rising because
hot water @image2 is poured into it. The pour follows the curve of a
matcha whisk @image3 moving in slow circles.
Camera: top-down 90mm macro, slow zoom-out over 8 seconds.
Light: morning window light, soft directional from screen-left.
Style: minimalist product photography, Hasselblad X2D color science.
@audio: @audio1
在此示例中:
- @image1 绑定到”陶瓷茶杯” —— 使用其颜色、形状和纹理。
- @image2 绑定到”热水” —— 影响蒸汽行为和流体动力学。
- @image3 绑定到”抹茶刷” —— 锁定道具身份。
- @audio1 提供环境音景(搅拌声、注水声、室内底噪)。
在我们的测试中,多模态提示词能显著减少 产品展示 和 角色一致性 工作流的漂移问题。关于角色制作,另请参阅我们的 Ref2V 参考转视频指南。
30+ 分类提示模板
以下 12 个模板分为四大类。每个模板均经过生产验证,可直接复制使用。
电影叙事
1. 悬疑开场
A man in a charcoal trench coat steps out of a black sedan car at noon,
the door closing behind him because he has reached his target address.
The scene is a rain-soaked Tokyo side street at midnight, neon reflecting
in standing water. Camera: wide-to-medium push-in, 40mm anamorphic,
handheld micro-shake. Light: cold neon teal and magenta, sodium-vapor
overhead. Style: neo-noir, grain, Blade Runner 2049 color palette.
Sound: rain on asphalt, distant traffic, footsteps echoing.
2. 站台告别
A woman in a yellow dress embraces her partner on a train platform,
tears falling because the departure board shows the train leaving.
The scene is a 1970s European railway terminal, fog rolling in through
the open doors of a stationary train. Camera: slow track left, 50mm
prime, shallow depth of field, focus racking from the woman to the
departure board. Light: tungsten warm from platform lamps, cool blue
exterior. Style: period cinematic, Kodak Vision3 200T emulation.
Sound: steam hissing, muffled announcement, distant whistle.
3. 发现
An archaeologist brushes dust from a carved stone tablet in a dimly
lit chamber, the inscription glowing faintly because her torch
passes over the carved grooves. The scene is an Egyptian tomb
interior, hieroglyphs covering walls. Camera: extreme close-up on
the brush, slow tilt up to reveal the chamber. Light: warm torch
practical, deep shadow falloff. Style: cinematic realism, Roger
Deakins-style motivated sources. Sound: dust falling, breath, distant
wind down the corridor.
产品展示
4. 奢华腕表揭幕
A stainless-steel chronograph watch rests on black marble, the second
hand ticking because the movement has been wound. The scene is a
minimalist studio set with a single gradient backdrop. Camera:
macro 100mm, slow orbit 180°, focus stacking. Light: hard top-key,
soft fill from screen-left, sharp specular highlight on the crystal.
Style: high-end commercial photography, Phase One IQ4 detail.
Sound: quiet room tone, subtle mechanical tick.
5. 护肤品瓶
A frosted-glass serum bottle sits among scattered flower petals, the
cap lifting because the product has just been opened. The scene is
a marble vanity with morning window light. Camera: 85mm, slow tilt
from base to dropper, focus pull. Light: soft directional window
with bounce fill. Style: editorial beauty, neutral peach palette.
Sound: gentle cap release, water dripping.
6. 跑道上的运动鞋
A neon-accent running shoe plants firmly on a wet track surface, water
displacing outward because the runner has just landed a stride.
The scene is a rainy evening training session at an empty stadium.
Camera: ground-level 24mm, slow dolly-right tracking the footfall.
Style: sports-commercial, Sony Venice X-OCN, high-contrast color.
Sound: sneaker impact, breath, distant crowd murmur.
角色驱动
7. 主角登场
A female warrior in bronze armor steps into a sunlit temple courtyard,
her hand releasing the strap because she has reached a place of safety.
The scene is a Greco-Roman ruin at midday, marble columns casting long
shadows. Camera: full-body wide, slow push-in, 35mm. Light: high noon
sun with dappled shade. Style: epic cinematic, desaturated olive grade.
Sound: leather creak, sandalled footfall on stone, distant wind.
8. 静谧时刻
An elderly man in a cardigan sits in a worn armchair, his hand
releasing the book because he has fallen asleep mid-sentence.
The scene is a book-lined study at dusk, a single reading lamp
glowing. Camera: medium shot, static with gentle breathing motion,
85mm. Light: warm tungsten practical, cool ambient window. Style:
intimate drama, deep shadow falloff. Sound: clock ticking, page
settling, soft breath.
9. 风格化肖像
A young woman with silver hair @image1 turns her head because a
bird has landed on the windowsill beside her. The scene is a
loft apartment with afternoon side-light. Camera: close-up, 85mm,
slow parallax drift. Style: cinematic portrait, Fujifilm X-T5
film emulation, pastel palette. Sound: birdsong, distant city
hum, fabric rustle.
抽象 / 艺术
10. 液态色彩研究
Two streams of magenta and cyan ink collide in clear water, the
resulting turbulence forming fractal patterns because the flows
have opposing viscosities. The scene is a controlled macro tank
with black velvet background. Camera: macro 100mm, static, 4K
120fps for slow-motion playback. Light: dual soft side-lights,
each color-matched to its ink. Style: experimental, high-speed
photography aesthetic. Sound: muffled fluid dynamics, sub-bass
rumble.
11. 生成式建筑
A crystalline cathedral grows upward from a flat plain, spires
branching outward because the algorithm has reached its
terminal height parameter. The scene is a surreal dreamscape
at twilight, pastel sky gradient. Camera: wide-angle 16mm,
slow vertical tilt-up following the growth. Style: surreal
digital art, Inception-inspired geometry. Sound: crystalline
chimes, low drone, ascending choir.
12. 粒子编舞
A swarm of golden particles spirals upward from a candle flame,
the pattern breaking apart because a door in the background has shifted
the room's pressure. Camera: 85mm, static, 30 seconds at native 4K.
Light: single warm practical, deep contrast. Sound: white guiter
drone, single bowed glass note.
Additional (softer):
A slow string of piano keys plucks softly as the swarm reforms into
the outline of a human figure, dissolving because the door has
shut fully. Style: abstract artistic, IMAX-documentary aesthetic.
如需针对商业场景优化的更多模板,请参阅我们的 Wan 3.0 商业广告指南。
负向提示词
Seedance 2.5 对负向提示词响应良好 —— 即明确排除不期望出现的伪影。负向提示词在 API 中置于单独字段,或在网页 UI 中以 --no 前缀。
推荐的负向提示词基线
--no text overlays, watermarks, on-screen captions, frame numbers,
logo artifacts, lens distortion, chromatic aberration, motion blur,
flicker, jitter, morphing faces, extra limbs, extra digits, deformed
anatomy, uncanny valley skin, plastic skin, oversaturation,
harsh contrast, blown highlights, crushed shadows, banding,
noise, grain artifacts, duplicate frames, identity drift,
costume change, scene change
分类特定的负向提示词
| 类别 | 补充排除项 |
|---|---|
| 照片写实 | “cartoon, anime, illustration, painting, CGI, 3D render, stylization” |
| 时代/历史 | “modern clothing, smartphone, electric lights, cars, anachronisms” |
| 产品 | “human figures, faces, hands, busy backgrounds, cluttered set” |
| 角色 | “costume change, hair color change, eye color change, age shift, body morph” |
在我们的测试中,未经优化的 Seedance 2.5 输出最常见的伪影是 身份漂移(角色特征在片段中途发生变化)和 高光过曝(4K 动态范围被截断的过曝区域)。将”identity drift”和”blown highlights”加入负向提示列表,在我们的编辑评审中可有效减少这两类问题。
进阶:时间戳控制与局部编辑
Seedance 2.5 引入了两项生产团队依赖的精确功能:1 秒时间戳分段 和生成帧的 局部编辑。
时间戳分段
较长的提示可以使用 [0s-5s]、[5s-12s]、[12s-30s] 标记法显式划分为计时段落。模型将每个段落规划为关联镜头,而非即兴衔接转场。
[0s-8s] A woman in a vintage coat walks slowly through a snowy
park avenue, breath visible because the temperature is below
zero. Camera: wide tracking shot, 35mm.
[8s-18s] She pauses at a fountain, her hand reaching out to
touch the frozen water because she has recognized the sculpture.
Camera: medium push-in, 50mm, slow dolly.
[18s-30s] She turns and walks away, leaving a single set of
footprints in the snow. Camera: wide static, snow falling.
每个段落作为独立镜头进行规划,模型通过主体、光照和色彩的连续性将它们相连。在我们的测试中,与单一整体提示词相比,这种分段方式能大幅减少场景切换伪影。
局部编辑
API 支持 局部编辑 —— 重新生成特定的 1 秒窗口,而无需重新生成整个片段。生成后,用户可以标记有问题的段落(例如 [12s-14s]),并仅针对该段落提交修正提示,同时保留片段其余部分的帧。这对于以下生产工作流至关重要:30 秒片段中大部分内容良好,仅需修正少数几秒。
作为对比,Veo 3.1 的原生音频工作流处理方式有所不同 —— 请参阅我们的 Veo 3.1 原生音频 4K 指南。
常见问题
1. Seedance 2.5 的最佳提示长度是多少?
在我们的测试中,80 至 180 词 之间的提示能产出最稳定的结果。过短的提示留有太多随机性;过长的提示(300+ 词)会因注意力衰减导致模型丢失提示后半部分的元素。使用七部分公式并保持在范围内。
2. Seedance 2.5 能否在一条提示中生成多镜头叙事?
可以。使用 时间戳分段([0s-8s]、[8s-18s])在单条提示中规划独立镜头,并将每个段落视为关联的节拍(new.qq.com)。
3. 一条提示中最多可以包含多少张参考图像和视频?
最多 30 张图像、10 个视频片段 和 10 个音频文件 —— 合计上限为 50 个多模态素材。通过 @image1、@video1、@audio1 语法分别寻址(new.qq.com)。
4. Seedance 2.5 是否生成原生音频?
Seedance 2.5 根据提示文本生成同步的环境音轨,并通过 @audio 标签接受最多 10 个音频文件参考。如需在原生 4K 下完全生成对话、音乐和牲畜音效,并配合同步唇形动作,请与 Veo 3.1 对比 —— 见我们的 Veo 3.1 指南。
5. Seedance 2.5 如何处理角色在不同镜头间的一致性?
为角色面部使用 @image1 参考,并在每个时间戳段落中引用同一 @image1 标签。模型通过将视觉参考绑定到每段中最近的名词来保持身份一致性。关于此工作流的深入介绍,请参阅我们的 Ref2V 指南 和 Kling 3.0 角色一致性指南。
6. Seedance 2.5 支持哪些宽高比?
原生支持 16:9、9:16、1:1、4:3 和 3:4,分辨率涵盖原生 4K(2160p)和 1080p。在提示前以 [16:9] 或 [9:16] 指定宽高比。
7. Seedance 2.5 API 是否已公开发布?
已发布 —— API 于 2025 年 8 月 上线(chinanews.com.cn),可通过 ByteDance 的火山引擎及合作伙伴平台使用。
8. Seedance 2.5 提示词与 Veo 3.1 或 Kling 3.0 提示词相比如何?
每个模型偏好的提示结构不同。Veo 3.1 偏好更长、更对话式的提示,并附带明确的音频描述。Kling 3.0 偏好紧凑、以视觉为主导、镜头指令强的提示。如需横向对比及一套跨模型通用的提示框架,请参阅我们的 2026 年最佳 AI 视频提示词大师课。
结论
截至 2026 年 10 月,Seedance 2.5 是功能最强大的长篇文本转视频模型,其提示结构也比前代更为严谨。使用 [主体] + [动作] + [场景] + [镜头] + [光线] + [风格] + [声音] 公式,借助因果连接词保持连贯性,并在需要稳定身份、运动或声音时使用 @image/@video/@audio 标签。
如需阅读 videosprompt.org 库中的相关文章,建议从以下开始:
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章