AI 线稿到照片级视频提示词:ControlNet、Runway 与 Seedance 实战工作流
TL;DR
- 本工作流通过三阶段将手绘或 AI 生成的线稿转换为照片级动态画面:ControlNet 线稿转图像、静态图像精修、图像转视频(I2V)。
- 使用 ControlNet 的
control_v11p_sd15_lineart.pth,控制权重 0.85–0.92,可显著优于更低权重保留线条结构;权重过低会使线条被模糊化为噪声。 - Runway Gen-3、Seedance 2.5 与 Kling I2V 是本流程效果最强的视频后端;对明确禁止柔焦与光晕的提示词反应各异。
- 本文 12 个提示词模板按使用场景分组(角色、环境、产品、氛围),以围栏代码块呈现,可直接粘贴至 Stable Diffusion WebUI 或 ComfyUI。
- 边缘光晕、角色变形、颜色溢出等常见失败模式,掌握各帧观察要点后均可诊断修复。
- 对分镜师和动画师,本工作流可替代手关键帧动画预览;对产品和概念艺术家,则大幅缩短从草图到客户成片的路径。
本工作流能做什么——适用人群
若你曾眼睁睁看着 Stable Diffusion 抹掉草图中的干净轮廓线,或看着 Runway 把精心绘制的线稿角色变成蜡质色块,这套工作流正是为你准备的。
本流程以线稿图像(手绘分镜面板、漫画描线扫描稿、Cleanroom 矢量导出图或 AI 生成的纯线稿帧)为输入,输出短幅照片级视频片段,且原始线条结构在剪影、光照和运动弧线中清晰可见。它并非”线稿加滤镜”,而是三阶段流水线,在通常会破坏线条的扩散、超分和时序采样中显式保留线条完整性。
三类核心受众:
- 分镜师与前期视觉总监,需对分镜面板快速进行照片级渲染,向导演传达意图。
- 独立动画师与漫画创作者,希望在投入全帧关键帧动画前获得动态预览。
- 产品与概念艺术家,需向客户展示草图阶段的工业设计或氛围帧如何呈现为动态场景。
流程并非一键完成:静态图像超分 → 静态图像精修 → 视频动态三步。跳过阶段是线条丢失的最常见原因。
三阶段流水线
第一阶段——线稿到静态图像(ControlNet)
ControlNet 是将扩散模型输出锚定在线稿源上的唯一现实路径。保持线条同时避免结果像涂色书的配置如下:
| 参数 | 数值 | 原因 |
|---|---|---|
| 模型 | control_v11p_sd15_lineart.pth |
专为干净纯线稿输入训练 |
| 预处理器 | lineart_standard |
在 ControlNet 读取前将任意图像转为干净线稿 |
| 控制权重 | 0.85–0.92 | 低于 0.8 时扩散会”在线条周围填充”;高于 0.95 时变成扁平插画 |
| 起始步数 | 0 | ControlNet 必须从第一个去噪步骤开始引导 |
| 结束步数 | 1.0(全程) | 过早放手会导致后期细化阶段漂移 |
| CFG 尺度 | 7–9 | 高于 10 会压缩照片级光照通道 |
| 基础模型 | Realistic Vision v5.1、Juggernaut XL 或 SDXL 基础模型 + 写实 LoRA | 写实型 checkpoint 对线条约束的处理比动漫型更忠实 |
lineart_standard 几乎适用于所有输入。lineart_coarse 仅适用于粗糙姿态速写、希望 ControlNet 推断结构的场景;lineart_realistic 倾向过度平滑,反而会重新引入本应抑制的照片纹理。
第二阶段——静态图像精修
第一轮输出的图像通常难以交付。第二轮是一轮简短修复:
- 以 1.5x 比例重绘手部和面部。ControlNet 能保留手部*形状*,却无法保证*正确性*。对每只手使用 512x512 的局部重绘遮罩可修复最常见的失败。
- 以低强度 img2img(去噪强度 0.25–0.35)调整光照方向。在此决定光线来自镜头左方还是右方。
- 以 LUT 或基于参考照片的最终 ControlNet 通道进行色彩分级。若照片级阶段漂向扁平日光白平衡,色彩参考可将其重新锚定。
- 以写实型超分器(4x-UltraSharp、NMKD Siax 或 RealESRGAN_x4plus)进行超分。切勿使用动漫型超分器——它们会将线条间断重绘为墨痕,并引入第三阶段不得不对抗的光晕。
第二阶段输出一张已锁定的、画面就绪的静态图像,分辨率与目标视频一致(通常 1280x720 或 1920x1080)。该图像锁定前,勿启动第三阶段。
第三阶段——静态图像到视频(Runway、Seedance 或 Kling)
第三阶段将锁定静态图像转化为动态画面。以下三个后端值得测试:
| 后端 | 最适合 | 优势 | 注意事项 |
|---|---|---|---|
| Runway Gen-3 Alpha Turbo | 写实角色动态、电影感镜头 | 4 秒内时序连贯性最强 | 激进去噪可能弱化轮廓边缘——可通过超分前锐化抵消 |
| Seedance 2.5(I2V 模式) | 动漫/漫画线稿、风格化运动 | 跨帧线条保留度最佳 | 对人体皮肤的写实保真度略低于 Runway |
| Kling I2V v1.6 | 长镜头运动、环境动态 | 处理 5–10 秒片段时线条轮廓稳定 | 面部存在首帧漂移;需在第二阶段锁定面部 |
第三阶段提示词与第一阶段提示词有本质区别,必须显式保护前两阶段保留的线条。诸如”电影感、24fps、缓慢推近”之类的通用开头无法保护任何东西。
针对 Seedance,请参阅我们的 Seedance I2V 模式指南 和 Ref2V 参考图转视频工作流——二者均为本流水线的延伸。
线稿保留的 ControlNet 参数设置
提示词与 ControlNet 设置中的模糊措辞是线条退化的根本成因。”保留线稿”“保持我的绘画可见”“维持速写质量”等措辞对扩散模型毫无意义——它根本没有”你的绘画”这一类别。
| 按影响力排序的关键旋钮 | 参数 | 效果 | 线稿推荐值 |
|---|---|---|---|
| 控制权重 | ControlNet 对扩散的约束强度 | 0.85–0.92 | |
| 控制起始步 | ControlNet 开始生效的时机 | 0 | |
| 控制结束步 | ControlNet 停止生效的时机 | 1.0 | |
| 预处理器分辨率 | 线条提取的分辨率 | 与输入图像一致;勿下采样 | |
| CFG 尺度 | 提示词被遵循的严格程度 | 7–9 | |
| 去噪强度(img2img) | 图像被重写的程度 | 0.45–0.55 用于完整渲染,0.25–0.35 用于精修 |
若输出看起来是照片级照片、完全看不到原始线稿,说明控制权重过低或去噪强度过高。若看起来是扁平插画且颜色已被填涂,说明控制权重过高或使用了动漫型基础模型。
若需深入了解现代 AI 流水线中草图输入的处理,请参阅 aitopia 草图转图像代理概览、telectronichub 复古动漫风格迁移线条完整性指南、Firered 草图转图像工具文档 与 Style3D 草图转换博客——各篇针对同一问题的不同阶段。
线稿的动态提示词编写
第三阶段提示词采用防御性写法,包含需保留的内容并明确排除会破坏线条的内容。
始终包含——以下短语在 Runway、Seedance 和 Kling 上均具可观测效果:
- “preserve line integrity”
- “sharp edges, no soft focus”
- “clean contour silhouette”
- “line-anchored motion”
- “photoreal lighting on a line-defined form”
始终排除——以下短语可阻止最常见的 I2V 退化:
- “anti-aliased edges”
- “soft focus, dreamy”
- “bokeh background blur on the subject”
- “halos around edges”
- “color bleed across the contour”
- “JPEG artifacts, compression noise”
“排除”列表不可妥协。扩散视频模型默认采用柔焦,因电影镜头即如此成像;没有明确对冲指令,I2V 通道会在前 3–4 帧内将硬朗墨线柔化为光晕。
12 个提示词模板
以下每个模板均可直接粘贴到 Stable Diffusion WebUI 的提示词框(与上表 ControlNet 设置搭配使用),或 ComfyUI 的 CLIPTextEncode 节点。第三阶段时,将”Animation prompt”部分复制到 Runway、Seedance 或 Kling。
角色动画(3)
Template 1 — Hero portrait, line-art warrior
Stage 1 prompt:
photoreal cinematic portrait of a young female warrior in bronze armor,
clean contour silhouette, sharp edges no soft focus, preserve line integrity,
shallow depth of field, golden hour rim light, 85mm lens, 8k uhd,
masterpiece quality, line-anchored detail
Negative: anti-aliased edges, soft focus, halos, color bleed, anime,
cartoon, illustration, jpeg artifacts
Stage 3 animation prompt (Runway / Seedance):
subtle wind moves hair across face, eyes blink once, armor catches
flickering torch light, camera slow push-in, preserve line integrity,
sharp edges no soft focus, clean contour silhouette, photoreal lighting
on line-defined form, 24fps cinematic motion
Template 2 — Manga panel to motion
Stage 1 prompt:
photoreal still of a teenaged boy sprinting through a rainy tokyo alley,
neon reflections on wet pavement, motion blur on background, sharp edges
on subject, preserve line integrity, clean contour silhouette,
35mm anamorphic look, cinematic color grade, 8k detail
Negative: anime, cel shading, flat color, halos, soft focus,
anti-aliased, jpeg artifacts
Stage 3 animation prompt:
character runs toward viewer, rain falls in streaks, neon flickers,
camera tracking backward at character speed, preserve line integrity,
sharp edges, no halos around character silhouette
Template 3 — Stylized mascot, product launch
Stage 1 prompt:
photoreal plush mascot character with oversized eyes, soft studio
lighting, sharp clean edges, preserve line integrity, line-anchored
form, product photography, white seamless backdrop, 8k product render
Negative: soft focus on subject, halos, anti-aliased, illustration,
anime, color bleed, painterly
Stage 3 animation prompt:
mascot turns head left, blinks, gives a small wave, eyes sparkle,
preserve line integrity, sharp edges no soft focus,
clean contour silhouette throughout motion
环境/分镜转场景(3)
Template 4 — Establishing shot
Stage 1 prompt:
photoreal mountain valley at sunrise, mist in the valley floor,
pine forest in foreground, sharp clean edges on tree silhouettes,
preserve line integrity, line-anchored composition, national geographic
style, 8k landscape, cinematic widescreen
Negative: soft focus, halos, painterly, illustration, anime,
color bleed, anti-aliased
Stage 3 animation prompt:
camera slow pan right, mist drifts left to right, sun rises,
birds cross frame, preserve line integrity, sharp edges on
tree silhouettes, no soft focus on foreground
Template 5 — Interior, noir lighting
Stage 1 prompt:
photoreal 1940s detective office interior, venetian blind shadows on
floor, single desk lamp light, sharp clean edges on furniture,
preserve line integrity, line-anchored detail, film noir cinematography,
8k interior, cinematic
Negative: anime, illustration, soft focus, halos, color bleed,
modern furniture, anti-aliased edges
Stage 3 animation prompt:
camera slow dolly in past desk lamp, venetian blind shadows shift
slowly as light source moves, cigarette smoke drifts, preserve line
integrity, sharp edges on all furniture silhouettes
Template 6 — Sci-fi corridor
Stage 1 prompt:
photoreal spaceship corridor, blinking overhead lights, steam from
floor grates, sharp clean edges on wall panels, preserve line integrity,
line-anchored industrial detail, sci-fi cinematography, 8k detail,
cinematic aspect ratio
Negative: soft focus, halos, illustration, anime, painterly,
anti-aliased, color bleed
Stage 3 animation prompt:
camera tracks forward down corridor, overhead lights flicker past,
steam rises from grates, emergency red light pulses, preserve line
integrity, sharp edges on wall panels, no halos around lights
产品/工业设计(3)
Template 7 — Industrial sketch to product render
Stage 1 prompt:
photoreal studio render of a matte black wireless earbud case,
subtle reflections on glass top, sharp clean edges on case body,
preserve line integrity, line-anchored form, product photography,
white seamless backdrop, 8k product render, advertisement quality
Negative: soft focus, halos around edges, illustration, anime,
color bleed, anti-aliased, painterly
Stage 3 animation prompt:
camera slow orbit around earbud case, lid opens smoothly, led
indicator pulses, preserve line integrity, sharp edges on case body,
no soft focus on product
Template 8 — Furniture concept
Stage 1 prompt:
photoreal mid-century modern lounge chair in walnut and tan leather,
soft window light from camera left, sharp clean edges on chair frame,
preserve line integrity, line-anchored furniture design, interior
design photography, 8k detail, architectural digest style
Negative: illustration, anime, soft focus, halos, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow orbit around chair, dust particles drift in window
light, leather catches highlights as light shifts, preserve line
integrity, sharp edges on chair frame throughout
Template 9 — Automotive sketch
Stage 1 prompt:
photoreal concept sports car in graphite gray, parked in empty
concrete showroom, single overhead spotlight, sharp clean edges on
body panels, preserve line integrity, line-anchored automotive design,
automotive photography, 8k detail, cinematic
Negative: soft focus, halos, illustration, anime, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow tracking shot from front quarter to side, headlamps
power on, reflections shift on body panels, preserve line integrity,
sharp edges on all body panels, no halos around headlamps
概念艺术/氛围(3)
Template 10 — Fantasy concept
Stage 1 prompt:
photoreal ancient stone temple in jungle, vines on crumbling pillars,
god rays through canopy, sharp clean edges on stonework, preserve line
integrity, line-anchored fantasy concept art, cinematic color grade,
8k environment, national geographic style
Negative: anime, illustration, soft focus, halos, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow crane up from ground level, god rays shift, vines sway
in breeze, a bird flies across frame, preserve line integrity,
sharp edges on stonework throughout
Template 11 — Horror mood
Stage 1 prompt:
photoreal abandoned hospital corridor, peeling paint on walls,
flickering fluorescent light, sharp clean edges on door frames,
preserve line integrity, line-anchored horror cinematography, desaturated
color grade, 8k interior, cinematic
Negative: soft focus, halos, illustration, anime, color bleed,
painterly, anti-aliased
Stage 3 animation prompt:
camera slow push-in down corridor, fluorescent light flickers,
a door creaks open at the end of the hall, preserve line integrity,
sharp edges on door frames, no soft focus on walls
Template 12 — Moody portrait
Stage 1 prompt:
photoreal close-up portrait of an older man with grey beard,
rain-streaked window behind him, single soft key light from camera
right, sharp clean edges on facial features, preserve line integrity,
line-anchored portrait photography, 8k detail, cinematic portrait
Negative: anime, illustration, soft focus on subject, halos,
color bleed, anti-aliased, painterly
Stage 3 animation prompt:
subject slowly turns head toward camera, eyes narrow slightly,
rain streaks shift on window behind, preserve line integrity,
sharp edges on facial features, no halos around face silhouette
对角色密集型创作,角色一致性提示词 2026 指南 详述如何在本流程中跨片段保持同一角色。对多镜头分镜,Seedance 多参考图提示词 将此方法扩展到跨镜头一致的环境。
工作流的重要性与 AI 工具的失败之处
线稿到照片级视频流水线以三种可预测方式失败,每种可从输出帧诊断,各有特定修复方法。
失败一——边缘光晕
外观表现。 最终视频中主体轮廓周围出现明亮辉光,仿佛扩散模型沿原始线条”泄漏”出了光照。
失败原因。 ControlNet 控制权重处于 0.65–0.75 区间——足以锚定扩散,但不足以阻止后期去噪步骤用与线条不完全匹配的照片级光照重绘剪影,由此呈现为辉光。
修复方法。 将控制权重提高至 0.85–0.92,并在第一、第三阶段提示词中加入”no halos around edges”。若光晕仍存在,问题出在超分器——将通用型超分器替换为 4x-UltraSharp 或 RealESRGAN_x4plus 等写实型超分器。
失败二——角色变形
外观表现。 到第 24 帧时,角色下颌线已向左偏移两像素,或鼻梁变窄,或发型发生微妙变化。到第 72 帧时,角色已与第 0 帧明显不同。
失败原因。 I2V 模型将角色视为”场景中的物体”而非锚定的身份。这是扩散视频模型的已知弱点,也是 I2V 角色锁定提示词 存在的主要原因。
修复方法。 使用 Ref2V(Seedance)或显式以面部标志点、发丝轮廓和服装轮廓为锚点的角色锁定提示词。在第二阶段锁定首帧,不允许 I2V 模型重新生成首帧。将动态提示词限制为小幅度、缓慢的运动;大幅度镜头运动会引发更多变形。
失败三——颜色溢出
外观表现。 一个元素的颜色扩散到相邻区域——天空颜色染到角色头顶,草地颜色染到建筑底部,灯光颜色染到灯周围一米范围内所有物体。
失败原因。 CFG 尺度过高(高于 10),模型通过过度饱和光照通道进行补偿。或第二阶段的色彩分级 LUT 过于激进,被 I2V 通道进一步放大。
修复方法。 将第一阶段 CFG 尺度降至 7–9,第二阶段使用更轻量级的色彩分级。在第一、第三阶段提示词中加入”no color bleed across contour”。若溢出集中在皮肤上,则基础模型选择有误——将动漫型 checkpoint 替换为写实型 checkpoint。
上述三种失败模式占”线条消失”问题的约 90%。其余 10% 可在 Runway 或 Seedling 预览中逐帧查看输出、定位线条首次退化的帧、再反向追溯原因(第一、第二还是第三阶段)来诊断。
常见问题
能否跳过 ControlNet,直接将原始线稿作为 Seedance 参考图像?
技术上可以,但将丧失本流程的大部分线条完整性。Seedance(以及 Runway、Kling)将参考图像视为*风格与构图线索*,而非*结构约束*。若第一阶段缺少 ControlNet 约束,照片级扩散会自由重绘剪影,到达第三阶段时线稿便不可见。ControlNet 正是将参考图像从*建议*变为*约束*的关键。若因时间原因需跳过 ControlNet,请接受输出呈现为”受线稿启发的照片级照片”,而非”保留线条结构的照片级渲染”。
漫画风格线稿的最佳 ControlNet 模型是什么?
针对漫画,应将 control_v11p_sd15_lineart.pth 与动漫-写实混合型 checkpoint(如 Counterfeit v3 或 MeinaMix)搭配,权重 0.78–0.85。漫画线宽变化大于西方描线,权重超过 0.85 会将细线重绘为粗线。若漫画带有网点效果,请先运行 lineart_standard,再通过快速亮度/对比度调整去除网点——否则网纹会作为”线条”进入约束条件。
能否使用 SDXL 或 Flux 替代 SD 1.5 完成本工作流?
可以,但有限制。SDXL 在光照处理上比 SD 1.5 更写实,但其 ControlNet 线稿模型仍不够成熟。对 SDXL,应使用 sai_xl_lineart 预处理器搭配 controlnet-sdxl-lineart,权重 0.75–0.85——比 SD 1.5 更低,因 SDXL 本身更忠实。截至 2026 年 10 月,Flux 尚无原生 ControlNet 支持;对 Flux 而言,最接近的方案是通过 IP-Adapter 进行参考图像约束,但其线条保留能力弱于 ControlNet。
最终视频片段应持续多长时间?
对 Runway Gen-3,4 秒是线条保留最佳时长——时序连贯性可保持约 4 秒,超过 6 秒后开始退化。对 Seedance 2.5,可通过 Ref2V 锚定延长至 8–10 秒。对 Kling I2V,5 秒是角色变形可见前的实际边界。若需更长片段,建议渲染两段并在后期拼接,而非让 I2V 模型超出其连贯性窗口。
是否需要高性能 GPU?
第一和第二阶段,12GB 显存 GPU(RTX 4070 及以上)可轻松应对 SD 1.5 + ControlNet。SDXL 需 16GB 作为最低配置。第三阶段,I2V 运算发生在 Runway、Seedance 或 Kling 的云端——本地 GPU 无关紧要,但超过 4 秒的片段长度需付费订阅。
本流水线能否用于手绘动画帧而非单张草图?
可以,但工作流需调整。对手绘动画,需将每个关键帧分别通过第一阶段处理,在第二阶段逐帧精修,然后以多参考图像的形式输入第三阶段。这就是 Seedance 多参考图扩展 工作流——相比单帧流水线可产出更流畅的角色动画,但要求各关键帧间角色线宽一致。
为何输出看起来是扁平插画而非照片级效果?
问题出在基础模型。动漫型 checkpoint(Anything v5、Counterfeit、MeinaMix)即使在强 ControlNet 权重下也会输出扁平插画,因 checkpoint 本身就是在扁平美学上训练的。若要获得照片级效果,请使用 Realistic Vision、Juggernaut 或 epiCRealism 作基础模型。即使 ControlNet 设置与提示词完美无缺,若 checkpoint 偏向插画风格,输出仍将是插画风格。
结语
线稿到照片级视频流水线是少有能让结果真正兑现”纸笔速写”承诺的 AI 工作流。草图在剪影、光照和运动弧线中保持可见。分镜师可在数小时内获得导演就绪的预览;产品设计师可在一下午获得照片级转台展示;概念艺术家则可获得客户所期待的具有线条完整性的氛围片。
代价是纪律性——三个阶段各有特定设置与失败方式。上述 12 个模板只是起点——一旦理解每个短语出现在提示词中的*原因*,你便能编写自己的模板。
若你是视频阶段新手,建议从 Seedance I2V 模式指南 与 角色锁定提示词参考 开始。若要扩展到多参考图分镜,Seedance 多参考图提示词文章 是自然的进阶读物。关于跨片段角色一致性,请参阅 2026 角色一致性提示词指南。
随着 Runway、Seedance 与 Kling 不断推出新版本,流水线将持续演进。但核心原则——以 ControlNet 保证结构、以精修保证质量、以防御性提示词保证线条完整性——不变。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章