AI Line Art to Photoreal Video Prompts: A Practical Workflow with ControlNet, Runway, and Seedance
TL;DR
- The workflow converts hand-drawn or AI-generated line art into photoreal motion clips through three stages: ControlNet line-to-image, still-image refinement, and image-to-video (I2V) generation.
- ControlNet’s
control_v11p_sd15_lineart.pthpreprocessor with control weight 0.85-0.92 preserves line integrity better than lower weights, which blur strokes into noise. - Runway Gen-3, Seedance 2.5, and Kling I2V are the strongest video back-ends for this pipeline; each responds differently to prompts that explicitly forbid soft focus and halos.
- The 12 prompt templates in this article are grouped by use case (character, environment, product, mood) and ship in fenced blocks ready to paste into Stable Diffusion WebUI or ComfyUI.
- Common failure modes — edge halos, character morphing, color bleed — are diagnosable and fixable once you know what to look for in each frame.
- For storyboarders and animators, this workflow replaces hand-keyed animation previews; for product and concept artists, it shortens the path from sketch to client-ready clip.
What this workflow delivers — and who it’s for
If you have ever watched a Stable Diffusion render quietly erase the clean contour lines in your sketch, or watched Runway turn a beautiful line-art character into a waxy blob, this workflow is for you.
The pipeline takes a line-art image — a hand-drawn storyboard panel, a manga inking scan, a Cleanroom block vector export, or an AI-generated line-only frame — and produces a short photoreal video clip where the original line structure is still visible in the silhouette, lighting, and motion arcs. It is not “line art + filter.” It is a three-stage pipeline that explicitly carries line integrity across the diffusion, upscaling, and temporal sampling steps that normally destroy it.
Three audiences rely on this:
- Storyboard artists and pre-vis supervisors who need a quick photoreal pass over a board panel to communicate intent to a director.
- Indie animators and manga creators who want a motion preview before committing to full-frame key animation.
- Product and concept artists who need to show a client how a sketched industrial design or mood frame reads as a living scene.
The pipeline is not a single click. It is still image upscale → still image refinement → video motion. Skipping stages is the most common reason line work disappears.
The three-stage pipeline
Stage 1 — Line art to still image (ControlNet)
ControlNet is the only realistic way to keep a diffusion model’s output anchored to a line-art source. The specific configuration that holds lines without making the result look like a coloring book is:
| Parameter | Value | Why |
|---|---|---|
| Model | control_v11p_sd15_lineart.pth |
Trained specifically on clean line-only inputs |
| Preprocessor | lineart_standard |
Converts any image to clean line art before ControlNet reads it |
| Control weight | 0.85–0.92 | Below 0.8 the diffusion “fills around” the lines; above 0.95 the result becomes a flat illustration |
| Starting step | 0 | ControlNet must guide from the very first denoising step |
| Ending step | 1.0 (full) | Letting go early causes late-stage refinements to drift |
| CFG scale | 7–9 | Higher than 10 begins to crush the photoreal lighting pass |
| Base model | Realistic Vision v5.1, Juggernaut XL, or SDXL base + realistic LoRA | Photoreal checkpoints handle line conditioning more honestly than anime checkpoints |
lineart_standard is the right preprocessor for almost every input. lineart_coarse is useful only for rough gesture sketches where you want ControlNet to infer structure; lineart_realistic tends to over-smooth and reintroduce the photo texture you are trying to control against.
Stage 2 — Still image refinement
The first-pass image is rarely shippable. The second pass is a short, targeted round of fixes:
- Inpaint hands and faces at 1.5x scale. ControlNet preserves the shape of hands; it does not preserve the correctness of hands. A 512x512 inpaint mask over each hand fixes the most common failure.
- Refine lighting direction with a low-strength img2img pass (denoising strength 0.25-0.35). This is where you decide whether the light is coming from camera-left or camera-right.
- Color grade with a LUT or a final ControlNet pass using a reference photo. If the photoreal stage has drifted toward a flat daylight white balance, a color reference re-anchors it.
- Upscale with a photoreal upscaler (4x-UltraSharp, NMKD Siax, or RealESRGAN_x4plus). Avoid anime upscalers — they will redraw line gaps as ink strokes and introduce the very halos Stage 3 then has to fight.
The output of Stage 2 is a single, frame-ready still image at your target video resolution (typically 1280x720 or 1920x1080). Do not start Stage 3 until this image is locked.
Stage 3 — Still image to video (Runway, Seedance, or Kling)
The third stage takes the locked still and produces motion. Three back-ends are worth testing for this pipeline:
| Back-end | Best for | Strength | Watch out for |
|---|---|---|---|
| Runway Gen-3 Alpha Turbo | Photoreal character motion, cinematic camera moves | Strongest temporal coherence at 4s | Aggressive denoising can soften contour edges — counter with a sharp pre-upscale |
| Seedance 2.5 (I2V mode) | Anime/manga line art, stylized motion | Best line preservation across frames | Slightly lower photoreal fidelity than Runway for human skin |
| Kling I2V v1.6 | Long camera moves, environmental motion | Handles 5-10s clips with consistent line silhouettes | First-frame drift on faces; lock the face in Stage 2 |
The animation prompt you write for Stage 3 is a different kind of prompt than the Stage 1 prompt. It must explicitly protect the lines you just spent two stages preserving. Generic starters like “cinematic, 24fps, slow dolly in” will not protect anything.
For Seedance specifically, see our Seedance I2V mode guide and the Ref2V reference-to-video workflow — both extend this exact pipeline.
ControlNet settings for line preservation
Vague terminology in prompts and ControlNet settings is the single biggest cause of line degradation. Words like “preserve the line art,” “keep my drawing visible,” or “maintain sketch quality” mean nothing to a diffusion model — it does not have a category for “your drawing.”
The knobs that matter, in order of impact | Parameter | Effect | Recommended for line art | |—|—|—| | Control weight | How strongly ControlNet conditions the diffusion | 0.85-0.92 | | Starting control step | When ControlNet starts guiding | 0 | | Ending control step | When ControlNet stops guiding | 1.0 | | Preprocessor resolution | Resolution at which the line is extracted | Match the input image; do not downsample | | CFG scale | How literally the prompt is followed | 7-9 | | Denoising strength (img2img) | How much the image is rewritten | 0.45-0.55 for full render, 0.25-0.35 for refinement |
If your output looks like a photoreal photo with no trace of the original line art, your control weight is too low or your denoising strength is too high. If it looks like a flat illustration with painted-in color, your control weight is too high or you are using an anime base model.
For deeper workflow context on how sketch inputs are handled in modern AI pipelines, see the aitopia sketch-to-image agent overview, the telectronichub linework-integrity guide for vintage anime style transfer, Firered’s sketch-to-image tool documentation, and the Style3D blog on sketch conversion — each addresses a different stage of the same problem.
The animation prompt for line art
Stage 3 prompts are written defensively. They include what to preserve and explicitly exclude what destroys lines.
Always include — these phrases have measurable effect across Runway, Seedance, and Kling:
- “preserve line integrity”
- “sharp edges, no soft focus”
- “clean contour silhouette”
- “line-anchored motion”
- “photoreal lighting on a line-defined form”
Always exclude — these phrases block the most common I2V degradations:
- “anti-aliased edges”
- “soft focus, dreamy”
- “bokeh background blur on the subject”
- “halos around edges”
- “color bleed across the contour”
- “JPEG artifacts, compression noise”
The “exclude” set is non-negotiable. Diffusion video models default to soft focus because that is how cinema lenses render; without an explicit counter-instruction, the I2V pass will soften your hard ink lines into halos within the first 3-4 frames.
12 prompt templates
Each template below is ready to paste into Stable Diffusion WebUI’s prompt field (paired with the ControlNet settings in the table above) or into ComfyUI’s CLIPTextEncode. For Stage 3, copy the “Animation prompt” into Runway, Seedance, or Kling.
Character animation (3)
Template 1 — Hero portrait, line-art warrior
Stage 1 prompt:
photoreal cinematic portrait of a young female warrior in bronze armor,
clean contour silhouette, sharp edges no soft focus, preserve line integrity,
shallow depth of field, golden hour rim light, 85mm lens, 8k uhd,
masterpiece quality, line-anchored detail
Negative: anti-aliased edges, soft focus, halos, color bleed, anime,
cartoon, illustration, jpeg artifacts
Stage 3 animation prompt (Runway / Seedance):
subtle wind moves hair across face, eyes blink once, armor catches
flickering torch light, camera slow push-in, preserve line integrity,
sharp edges no soft focus, clean contour silhouette, photoreal lighting
on line-defined form, 24fps cinematic motion
Template 2 — Manga panel to motion
Stage 1 prompt:
photoreal still of a teenaged boy sprinting through a rainy tokyo alley,
neon reflections on wet pavement, motion blur on background, sharp edges
on subject, preserve line integrity, clean contour silhouette,
35mm anamorphic look, cinematic color grade, 8k detail
Negative: anime, cel shading, flat color, halos, soft focus,
anti-aliased, jpeg artifacts
Stage 3 animation prompt:
character runs toward viewer, rain falls in streaks, neon flickers,
camera tracking backward at character speed, preserve line integrity,
sharp edges, no halos around character silhouette
Template 3 — Stylized mascot, product launch
Stage 1 prompt:
photoreal plush mascot character with oversized eyes, soft studio
lighting, sharp clean edges, preserve line integrity, line-anchored
form, product photography, white seamless backdrop, 8k product render
Negative: soft focus on subject, halos, anti-aliased, illustration,
anime, color bleed, painterly
Stage 3 animation prompt:
mascot turns head left, blinks, gives a small wave, eyes sparkle,
preserve line integrity, sharp edges no soft focus,
clean contour silhouette throughout motion
Environment / storyboard to scene (3)
Template 4 — Establishing shot
Stage 1 prompt:
photoreal mountain valley at sunrise, mist in the valley floor,
pine forest in foreground, sharp clean edges on tree silhouettes,
preserve line integrity, line-anchored composition, national geographic
style, 8k landscape, cinematic widescreen
Negative: soft focus, halos, painterly, illustration, anime,
color bleed, anti-aliased
Stage 3 animation prompt:
camera slow pan right, mist drifts left to right, sun rises,
birds cross frame, preserve line integrity, sharp edges on
tree silhouettes, no soft focus on foreground
Template 5 — Interior, noir lighting
Stage 1 prompt:
photoreal 1940s detective office interior, venetian blind shadows on
floor, single desk lamp light, sharp clean edges on furniture,
preserve line integrity, line-anchored detail, film noir cinematography,
8k interior, cinematic
Negative: anime, illustration, soft focus, halos, color bleed,
modern furniture, anti-aliased edges
Stage 3 animation prompt:
camera slow dolly in past desk lamp, venetian blind shadows shift
slowly as light source moves, cigarette smoke drifts, preserve line
integrity, sharp edges on all furniture silhouettes
Template 6 — Sci-fi corridor
Stage 1 prompt:
photoreal spaceship corridor, blinking overhead lights, steam from
floor grates, sharp clean edges on wall panels, preserve line integrity,
line-anchored industrial detail, sci-fi cinematography, 8k detail,
cinematic aspect ratio
Negative: soft focus, halos, illustration, anime, painterly,
anti-aliased, color bleed
Stage 3 animation prompt:
camera tracks forward down corridor, overhead lights flicker past,
steam rises from grates, emergency red light pulses, preserve line
integrity, sharp edges on wall panels, no halos around lights
Product / industrial design (3)
Template 7 — Industrial sketch to product render
Stage 1 prompt:
photoreal studio render of a matte black wireless earbud case,
subtle reflections on glass top, sharp clean edges on case body,
preserve line integrity, line-anchored form, product photography,
white seamless backdrop, 8k product render, advertisement quality
Negative: soft focus, halos around edges, illustration, anime,
color bleed, anti-aliased, painterly
Stage 3 animation prompt:
camera slow orbit around earbud case, lid opens smoothly, led
indicator pulses, preserve line integrity, sharp edges on case body,
no soft focus on product
Template 8 — Furniture concept
Stage 1 prompt:
photoreal mid-century modern lounge chair in walnut and tan leather,
soft window light from camera left, sharp clean edges on chair frame,
preserve line integrity, line-anchored furniture design, interior
design photography, 8k detail, architectural digest style
Negative: illustration, anime, soft focus, halos, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow orbit around chair, dust particles drift in window
light, leather catches highlights as light shifts, preserve line
integrity, sharp edges on chair frame throughout
Template 9 — Automotive sketch
Stage 1 prompt:
photoreal concept sports car in graphite gray, parked in empty
concrete showroom, single overhead spotlight, sharp clean edges on
body panels, preserve line integrity, line-anchored automotive design,
automotive photography, 8k detail, cinematic
Negative: soft focus, halos, illustration, anime, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow tracking shot from front quarter to side, headlamps
power on, reflections shift on body panels, preserve line integrity,
sharp edges on all body panels, no halos around headlamps
Concept art / mood (3)
Template 10 — Fantasy concept
Stage 1 prompt:
photoreal ancient stone temple in jungle, vines on crumbling pillars,
god rays through canopy, sharp clean edges on stonework, preserve line
integrity, line-anchored fantasy concept art, cinematic color grade,
8k environment, national geographic style
Negative: anime, illustration, soft focus, halos, color bleed,
anti-aliased, painterly
Stage 3 animation prompt:
camera slow crane up from ground level, god rays shift, vines sway
in breeze, a bird flies across frame, preserve line integrity,
sharp edges on stonework throughout
Template 11 — Horror mood
Stage 1 prompt:
photoreal abandoned hospital corridor, peeling paint on walls,
flickering fluorescent light, sharp clean edges on door frames,
preserve line integrity, line-anchored horror cinematography, desaturated
color grade, 8k interior, cinematic
Negative: soft focus, halos, illustration, anime, color bleed,
painterly, anti-aliased
Stage 3 animation prompt:
camera slow push-in down corridor, fluorescent light flickers,
a door creaks open at the end of the hall, preserve line integrity,
sharp edges on door frames, no soft focus on walls
Template 12 — Moody portrait
Stage 1 prompt:
photoreal close-up portrait of an older man with grey beard,
rain-streaked window behind him, single soft key light from camera
right, sharp clean edges on facial features, preserve line integrity,
line-anchored portrait photography, 8k detail, cinematic portrait
Negative: anime, illustration, soft focus on subject, halos,
color bleed, anti-aliased, painterly
Stage 3 animation prompt:
subject slowly turns head toward camera, eyes narrow slightly,
rain streaks shift on window behind, preserve line integrity,
sharp edges on facial features, no halos around face silhouette
For character-heavy work, the character-consistency prompts 2026 guide covers how to keep the same character across multiple clips in this pipeline. For multi-shot storyboards, Seedance multi-reference image prompts extend this approach to consistent environments across shots.
Why workflow matters and where AI tools fail
The line-art-to-photoreal-video pipeline fails in three predictable ways. Each is diagnosable from the output frames, and each has a specific fix.
Failure 1 — Edge halos
What it looks like. A bright glow appears around the contour of the subject in the final video, like the diffusion model “leaked” light along the original line.
Why it happens. Your ControlNet control weight is in the 0.65-0.75 range — strong enough to anchor the diffusion but weak enough that the late denoising steps redraw the silhouette with photoreal lighting that doesn’t quite match the line. The mismatch reads as a glow.
How to fix it. Raise control weight to 0.85-0.92 and add “no halos around edges” to both Stage 1 and Stage 3 prompts. If the halo persists, your upscaler is the culprit — switch from a general-purpose upscaler to a photoreal upscaler like 4x-UltraSharp or RealESRGAN_x4plus.
Failure 2 — Character morphing
What it looks like. By frame 24, your character’s jawline has shifted two pixels left, or their nose has narrowed, or their hairstyle has subtly changed. By frame 72, the character is recognizably different from frame 0.
Why it happens. Your I2V model is treating the character as a “thing in the scene” rather than an anchored identity. This is a known weakness of diffusion video models and is the main reason I2V character lock prompts exist.
How to fix it. Use Ref2V (Seedance) or character-lock prompts that explicitly call out facial landmarks, hair silhouette, and clothing silhouette as anchors. Lock the first frame in Stage 2 — do not let the I2V model regenerate the first frame. Reduce the motion prompt to small, slow motions; large camera moves trigger more morphing.
Failure 3 — Color bleed
What it looks like. The color of one element spreads into adjacent areas — the sky tints the top of the character’s head, the grass color tints the bottom of the building, the lamp light tints everything within a meter of the lamp.
Why it happens. Your CFG scale is too high (above 10) and the model is over-saturating the lighting pass to compensate. Alternatively, your Stage 2 color-grade LUT is too aggressive and the I2V pass is amplifying it.
How to fix it. Drop CFG scale to 7-9 in Stage 1 and use a lighter color grade in Stage 2. Add “no color bleed across contour” to both Stage 1 and Stage 3 prompts. If the bleed is on skin specifically, your Stage 1 base model is wrong — switch from an anime checkpoint to a photoreal checkpoint.
These three failure modes account for roughly 90 percent of “the lines disappeared” complaints. The remaining 10 percent are diagnosable by frame-stepping through the output in Runway or Seedling’s preview, finding the frame where the line first degrades, and working backward to identify whether the cause is Stage 1, Stage 2, or Stage 3.
FAQ
Can I skip ControlNet and just use the original line art as a Seedance reference image?
Technically yes, but you will lose most of the line integrity that motivated the pipeline. Seedance (and Runway, and Kling) treat reference images as style and composition hints, not as structural constraints. Without ControlNet conditioning in Stage 1, the photoreal diffusion in Stage 1 redraws the silhouette freely, and by the time you reach Stage 3 the line art is invisible. ControlNet is what makes the reference image a constraint rather than a suggestion. If you need to skip ControlNet for time reasons, accept that your output will look like “photoreal photo inspired by line art” rather than “photoreal render with preserved line structure.”
What is the best ControlNet model for manga-style line art?
For manga specifically, pair control_v11p_sd15_lineart.pth with an anime-realistic hybrid checkpoint (like Counterfeit v3 or MeinaMix) at control weight 0.78-0.85. Manga line weights vary more than western inking, so going above 0.85 will redraw thin lines as thick ones. If your manga is screen-toned, run lineart_standard first, then a quick brightness/contrast pass to drop the screentones before ControlNet reads the image — otherwise the screentones become “lines” in the conditioning.
Can I use SDXL or Flux instead of SD 1.5 for this workflow?
Yes, with caveats. SDXL handles lighting more realistically than SD 1.5 but its ControlNet line models are still less mature. For SDXL, use the sai_xl_lineart preprocessor paired with controlnet-sdxl-lineart at 0.75-0.85 — lower than SD 1.5 because SDXL is already more literal. Flux has no native ControlNet support as of October 2026; for Flux, the closest analog is reference-image conditioning through IP-Adapter, which gives weaker line preservation than ControlNet.
How long should the final video clip be?
For Runway Gen-3, 4 seconds is the sweet spot for line preservation — temporal coherence holds up to about 4s and degrades after 6s. For Seedance 2.5, you can extend to 8-10s with Ref2V anchoring. For Kling I2V, 5s is the practical limit before character morphing becomes visible. If you need a longer clip, render two clips and join them in post rather than asking the I2V model to extend past its coherence window.
Do I need a powerful GPU?
For Stage 1 and Stage 2, a 12GB VRAM GPU (RTX 4070 or above) handles SD 1.5 with ControlNet comfortably. For SDXL, 16GB is the realistic minimum. For Stage 3, the I2V work happens on Runway’s, Seedance’s, or Kling’s cloud — your GPU is irrelevant, but you do need a paid subscription for clip lengths above 4s.
Can this pipeline work with hand-drawn animation frames rather than a single sketch?
Yes, but the workflow changes. For hand-drawn animation, run each keyframe through Stage 1 separately, refine each in Stage 2, and feed them as multi-reference images into Stage 3. This is the multi-reference Seedance extension workflow — it produces smoother character animation than the single-image pipeline but requires consistent character line weights across your keyframes.
Why does my output look like a flat illustration instead of photoreal?
Your base model is the issue. Anime checkpoints (Anything v5, Counterfeit, MeinaMix) produce flat illustrations even with strong ControlNet weights because the checkpoint itself is trained on a flat aesthetic. For a photoreal result, use Realistic Vision, Juggernaut, or epiCRealism as your base model. Your ControlNet settings and prompts can be perfect, but if the checkpoint is illustration-trained, the output will be illustration-styled.
Conclusion
The line-art-to-photoreal-video pipeline is one of the few AI workflows where the result genuinely matches the sketch-on-paper promise. The sketch stays visible in the silhouette, the lighting, and the motion arc. Storyboarders get a director-ready preview in hours rather than days. Product designers get a photoreal turntable in an afternoon. Concept artists get a mood piece with the line integrity their clients expect.
The cost is discipline: three stages, each with specific settings, each failing in specific ways. The 12 templates above are starting points — once you understand why each phrase is in the prompt, you can write your own.
If you are new to the video stages, start with the Seedance I2V mode guide and the character-lock prompts reference. If you are extending this to multi-reference storyboards, the Seedance multi-reference image prompts article is the natural follow-up. For character consistency across clips, see the 2026 character-consistency prompts guide.
The pipeline will keep changing as Runway, Seedance, and Kling ship new versions. The principle — ControlNet for structure, refinement for quality, defensive prompting for line integrity — will not.
Reviewed by the videosprompt.org editorial team · October 2026
Share Article