VideosPrompt VideosPrompt

Image-to-Video Character Lock Prompts: Subject-Locked I2V Workflows

Author: VideosPrompt Date: 2026-10-07 14:11:42
Image-to-Video Character Lock Prompts: Subject-Locked I2V Workflows

TL;DR

  • I2V vs Ref2V: I2V starts from one still image and conditions motion on it. Ref2V starts from a subject extracted from reference photos. Both lock identity; the prompt grammars differ.
  • Character lock is a prompt-level contract: the prompt must state what to preserve (face, wardrobe, accessories, build) and what may change (pose, expression, camera, scene).
  • Lock strength is a continuum: soft allows wardrobe and pose drift; medium preserves face and clothing but allows re-poses; hard freezes everything except micro-expressions and physics-driven motion.
  • Starting image quality sets the ceiling: a sharp, front-facing, single-subject, neutrally posed, high-contrast source image reduces identity drift dramatically.
  • Models handle locking differently: Seedance 2.5 (reference image + motion video pair), Veo 3.1 (image-prompt + audio cue), Kling 3.0 (elements panel), Wan 3.0 (start-end frame pairing).
  • Twelve tested templates follow, grouped by soft / medium / hard / multi-character lock, each in a fenced text code block.
  • Lock-strength modifiers like "preserve identity exactly", "no face morphing", and "strict character continuity" sharpen preservation and must be stated explicitly.

Intro: What “character lock” actually means in an I2V pipeline

Image-to-video (I2V) models take a single still image and synthesize the motion that comes “next.” The default behavior of most I2V models is creative interpolation: they will re-light the face, re-shape the jawline, swap the jacket for a similar one, or age the subject if it makes the motion read better. Character lock is the set of prompt techniques — and starting-image choices — that tell the model “do not re-imagine this identity, just move it.”

Character lock in I2V differs from character lock in Ref2V (reference-to-video). I2V is image-conditioned motion: the model sees one still frame and continues it, locking the entire scene’s starting state pixel-for-pixel. Ref2V is identity-conditioned generation: the model sees reference photos and generates motion that may not match any single reference’s lighting, framing, or pose — locking identity only while ceding everything else to the prompt. See the companion guide on Seedance reference-to-video workflows and the sibling article AI character consistency prompts in 2026 for broader strategies.

Why character lock matters:

  1. Asset reusability. A locked character can be re-used across a campaign — social ads, explainers, training videos — without re-shooting.
  2. Narrative continuity. Recurring hosts, avatars, and brand mascots require the audience to read the same character across episodes. Drift breaks immersion within seconds.
  3. Lip-sync integration. Compositing a generated video with a recorded voiceover exposes facial drift between still and motion as an uncanny mismatch.

Character lock is a learnable, promptable skill, but probabilistic: the same prompt behaves differently across Seedance 2.5, Veo 3.1, Kling 3.0, and Wan 3.0. This article covers universal principles, per-model mechanics, and twelve ready-to-paste templates.


The I2V character-lock landscape: how each major model handles it

Each flagship model exposes character locking through a different mechanism, but they converge on one contract: a still (or set of stills) defines identity, the prompt defines everything else, the model negotiates motion.

Model Lock mechanism Primary input(s) Prompt role Lock ceiling
Seedance 2.5 Reference image + motion video pair Subject stills + motion clip Scene, lighting, action, style High; multi-reference can lock a 50-asset cast (see Seedance multi-reference image prompts)
Veo 3.1 Image-prompt with audio cue Starting still + optional audio Long descriptive prompt High for short clips; weaker through full turn
Kling 3.0 Elements panel Face, wardrobe, body stills Compact prompt + panel keywords High for face, medium for accessories
Wan 3.0 Start-end frame pairing Start + end frame Short motion prompt Highest at endpoints; weaker mid-clip

Seedance 2.5 treats the reference still as the identity anchor and a separate motion reference as the motion blueprint. The prompt covers scene, lighting, style — motion comes from the reference video. See the Seedance 2.5 prompts guide and the VidMuse deep dive.

Veo 3.1 is image-prompt centric. One still drives identity; the prompt encodes everything else, including audio cues (speech, ambient noise). Most language-first of the four.

Kling 3.0 uses a structured elements panel — face, clothing, optional body — fused by the model while the narrative prompt drives the scene. See the Kling 3.0 character consistency guide.

Wan 3.0 uses start-end frame pairing: you supply both first and last frame; the model interpolates. Lock strength is highest at endpoints and weakest mid-clip.

For universal templates, see the 2K video model prompt templates collection.


The character-lock prompt formula: six elements that always earn their place

A reliable I2V character-lock prompt has six parts. You can omit one or two in a soft lock, but at hard-lock strength every element must be present.

# Element What it does Example fragment
1 Source-image anchor Reaffirms the still as identity source "Using the provided still as the identity anchor"
2 Character token description Names what is locked: face shape, hair, build, skin tone, signature features "mid-30s East Asian woman, round face, short black bob, dark brown eyes"
3 Scene content Describes what the character is doing "walks through a crowded night market, glancing left and right"
4 Camera Defines framing, lens, movement "handheld medium close-up, 35mm, shallow depth of field"
5 Style Encodes visual treatment: film, animation, photography "cinematic, anamorphic, slightly desaturated, 2.39:1"
6 Consistency check Tells the model what not to do "no face morphing, no wardrobe change, preserve identity exactly"

The first element is the most important and the most often skipped. Many I2V failures trace back to a prompt that described the character beautifully but never told the model the still is non-negotiable. Element 6 is where you stack lock-strength modifiers.

A useful pattern, drawn from AI video prompt engineering best practices, is to put the identity description immediately after the source-image anchor and before the action. This matches how attention flows in current video diffusion models: identity is read first, then motion is conditioned on it.


Choosing the right starting image: what makes a “good lock source”

The starting image is the foundation of the entire lock. A poor source image cannot be rescued by a clever prompt; a good one makes even a mediocre prompt succeed. Five properties of a high-quality lock source:

  1. Sharp, well-lit face. Blurry, backlit, or occluded faces produce drift within the first 0.5 seconds. Eyes in focus, above 80px height.
  2. Frontal or three-quarter view. Profile shots hide half the identity signal. Front or three-quarter gives maximum anchors — both cheeks, both eyes, full brow.
  3. Single subject, separated from others. Multiple people force the model to guess the anchor, and the guess may shift mid-clip.
  4. High subject-background contrast. Blending into the background gives weak silhouette cues and increases edge artifacts.
  5. No occlusions on identity features. Sunglasses, hands over the mouth, deep hat shadows — each is a missed anchor.

A heuristic: if a casting assistant could pick this person out of a lineup of ten similar-looking extras, the source will lock well.

Resolution matters less than signal clarity — a clean 768px face still locks better than a noisy 4K still. Neutral lighting preserves skin tone fidelity; warm or cool casts in the source bake into the output.

For generated sources (text-to-image stills), apply these five properties before passing downstream.


Twelve prompt templates, organized by lock strength

Each template is a self-contained prompt for direct paste. Replace bracketed [placeholders] with your own values.

Soft lock (3 templates)

Soft locks allow the model significant creative latitude on pose, lighting, and wardrobe. Use when motion quality matters more than identity fidelity — B-roll, ambient shots, scene-setting cutaways.

Template SL-01 — Street B-roll, soft lock, model: universal
Using the provided still as the identity anchor.
Subject: [mid-30s subject, neutral pose, simple background] from the anchor still.
Action: walks away from camera down a tree-lined street, soft natural daylight.
Camera: handheld 50mm, slow push-in, subject centered in frame.
Style: cinematic, naturalistic color, 2.39:1 anamorphic feel.
Consistency: wardrobe may shift subtly with movement; preserve face shape and skin tone; allow natural lighting variation.
Template SL-02 — Ambient scene, soft lock, model: Seedance 2.5 / Veo 3.1
Reference still is the identity anchor. Motion video provides walk cycle.
Subject: [anchor subject, full body, neutral stance] from the still.
Action: enters a sunlit café, places a coffee cup on a table, smiles.
Camera: static medium shot at table height, 35mm, warm tungsten bounce.
Style: editorial lifestyle photography, soft grain, gentle vignette.
Consistency: identity preservation is important but motion realism takes priority; allow minor wardrobe drift toward similar palette; do not change hair color or facial structure.
Template SL-03 — Establishing shot, soft lock, model: Kling 3.0 / Wan 3.0
Use anchor still as subject identity. End frame is the still itself.
Subject: [anchor subject, three-quarter view, single subject frame] from the still.
Action: stands at a window in a high-rise apartment at dusk, city skyline behind, gently turning toward the camera.
Camera: locked-off wide shot, 24mm, deep focus, golden hour palette.
Style: cinematic still photography, slight film bleach, 16:9.
Consistency: preserve identity; minor pose drift toward end frame is acceptable; preserve skin tone and silhouette.

Medium lock (3 templates)

Medium locks preserve face and wardrobe but allow pose, gesture, and gaze changes. The most common lock strength for narrative content, talking-head explainers, and short-form social.

Template ML-01 — Talking head, medium lock, model: Veo 3.1
Anchor still is the identity source. Subject speaks the line below.
Subject: [anchor subject, sharp face, front view] from the still.
Action: speaks directly to camera with calm, measured delivery, line: "Welcome back to the series — today we look at three prompts that changed how I shoot video."
Camera: locked-off medium close-up, 50mm, eye level, soft key light at 45 degrees.
Style: clean studio look, neutral background, broadcast-safe color.
Consistency: preserve identity exactly; wardrobe must match the still including collar, lapel, and accessories; only the mouth, eyes, and head may move; no face morphing.
Template ML-02 — Walking interview, medium lock, model: Seedance 2.5
Anchor still defines subject identity. Motion reference is the supplied walk-and-talk clip.
Subject: [anchor subject, three-quarter view, walking pace] from the still.
Action: walks along a coastal path while speaking, gesturing occasionally with the right hand.
Camera: tracking alongside at shoulder height, 35mm, shallow depth of field.
Style: documentary cinematography, handheld feel, neutral grade.
Consistency: preserve identity exactly; wardrobe must match including hat, jacket, and bag strap; allow natural hand gestures and head turns; no face morphing; no wardrobe swap.
Template ML-03 — Re-pose sequence, medium lock, model: Kling 3.0
Anchor face and anchor wardrobe stills are the identity sources.
Subject: [anchor subject, two registered elements: face and wardrobe] from the elements panel.
Action: transitions through four poses over six seconds — front, three-quarter left, three-quarter right, and back to front — in a sunlit room.
Camera: slowly orbiting 360 around subject at 1m radius, 50mm, even exposure.
Style: clean editorial, soft shadow, light beige background.
Consistency: preserve identity exactly across all four poses; wardrobe must match in every pose including accessories; only head, neck, and arms may move; no face morphing; strict character continuity.

Hard lock (3 templates)

Hard locks freeze the subject down to small accessories — earrings, watch, moles, exact collar fold. The lock strength for series continuity, branded content, and any production where the audience has already met the character.

Template HL-01 — Series intro shot, hard lock, model: Veo 3.1
Anchor still is the absolute identity source. Subject speaks the line.
Subject: [anchor subject, exact front view, sharp face] from the still.
Action: a single calm breath, slight smile, then delivers the line: "Welcome to episode one."
Camera: locked-off medium close-up, 85mm, eye level, three-point lighting matching the still.
Style: cinematic, polished, brand-grade color.
Consistency: preserve identity exactly including every facial feature, the small mole near the left eyebrow, the earring in the right ear, the exact collar fold of the shirt; only mouth, eyes, and subtle head movement; no face morphing; no wardrobe change; strict character continuity.
Template HL-02 — Action with retained identity, hard lock, model: Seedance 2.5
Anchor still defines subject. Motion reference is a supplied action clip of the same subject.
Subject: [anchor subject, action pose, sharp face] from the still.
Action: climbs three rungs of a fire escape ladder while looking back over the shoulder toward lens.
Camera: low-angle handheld, 35mm, dramatic perspective, dusk light.
Consistency: preserve identity exactly including every facial feature, jacket color, jacket zipper position, watch on left wrist; allow full body motion including legs, arms, and torso; no face morphing; strict character continuity; preserve accessories exactly.
Template HL-03 — Profile with continuity, hard lock, model: Wan 3.0
Start frame: anchor still. End frame: anchor still rotated 90 degrees to subject's left.
Subject: [anchor subject, exact profile view from end frame].
Action: subject turns head slowly from right profile to left profile over four seconds.
Camera: locked-off close-up profile, 100mm, soft window light.
Style: minimalist editorial, neutral backdrop.
Consistency: preserve identity exactly including every facial feature and the visible side of the hair; allow only the rotational motion; no face morphing; strict character continuity; preserve accessories including earring and hair clip.

Identity-preserving cast (3 templates, multiple characters)

Multi-character locks are the hardest case because the model must lock two or more identities and keep them distinct. Describe each subject by distinguishing features rather than position alone — position-based references get confused when subjects cross.

Template MC-01 — Two-character dialogue, cast lock, model: Veo 3.1
Two anchor stills define identities: still A (left character), still B (right character).
Subject A: [anchor A, distinguishing features] seated left.
Subject B: [anchor B, distinguishing features] seated right.
Action: A speaks a line, then B responds; both maintain eye contact with each other, not the camera.
Camera: locked two-shot at 50mm, slight low angle, warm key from screen-left.
Style: cinematic dialogue scene, naturalistic color.
Consistency: preserve both identities exactly; do not let A's features drift toward B's or vice versa; preserve each character's wardrobe including all accessories; no face morphing; strict character continuity for both subjects.
Template MC-02 — Three-character group shot, cast lock, model: Seedance 2.5
Three anchor stills define identities. Motion reference is a supplied group action clip.
Subject A: [anchor A, distinctive wardrobe and features], positioned left.
Subject B: [anchor B, distinctive wardrobe and features], positioned center.
Subject C: [anchor C, distinctive wardrobe and features], positioned right.
Action: the trio walks together across an open plaza, B in the lead, C catching up.
Camera: wide tracking shot, 28mm, even daylight, deep focus.
Style: lifestyle commercial, polished, warm grade.
Consistency: preserve all three identities exactly across the entire motion; no face morphing on any subject; wardrobe must match including hats, bags, and visible logos; do not swap wardrobe between subjects; strict character continuity for the full cast.
Template MC-03 — Confrontation scene, cast lock, model: Kling 3.0
Two anchor face elements and two anchor wardrobe elements registered in the elements panel.
Subject A: [anchor A face element + anchor A wardrobe element].
Subject B: [anchor B face element + anchor B wardrobe element].
Action: A and B stand face to face in a narrow corridor; A leans in slightly while speaking.
Camera: handheld over-the-shoulder alternating close-ups, 50mm, cool fluorescent lighting.
Style: dramatic, high-contrast, desaturated.
Consistency: preserve both identities exactly; preserve both wardrobes exactly; do not let A's face drift into B's or vice versa; preserve all accessories including A's pendant and B's wristwatch; no face morphing; strict character continuity for both subjects.

Lock-strength modifiers: the verbal tokens that sharpen preservation

Lock strength depends on starting image, prompt content, and the exact words used to signal strictness. The modifiers below, ordered soft to hard, are the ones we have found most reliable across the four I2V models.

Modifier token Lock signal When to use
"allow natural lighting variation" Soft B-roll, ambient, atmospheric
"preserve face shape and skin tone" Soft-medium Lifestyle, documentary, editorial
"preserve identity" Medium Narrative, explainers, social
"preserve identity exactly" Medium-hard Series continuity, branded content
"no face morphing" Hard Facial features must remain stable
"no wardrobe swap" Hard Branded looks, recurring characters
"strict character continuity" Hard Multi-episode series, recurring hosts
"preserve accessories exactly" Hardest Watches, earrings, glasses, moles must remain visible
"no identity drift between frames" Hardest Benchmark / frame-level verification

A common mistake is using too many hard modifiers in a soft-lock prompt. The model can interpret "no face morphing" + "preserve identity exactly" + "strict character continuity" as conflicting when the scene requires re-expression (a smile transitioning to surprise, for example). For emotional arcs, dial back to "preserve identity".

Another common mistake is using no modifiers and expecting the model to infer preservation from context. The model infers motion priorities from action verbs, not preservation. Without explicit preservation language, motion wins.

Litmus test: read your prompt aloud and underline every word describing what the character does. If those outnumber your preservation tokens, the model treats motion as the goal and identity as decoration. Rebalance until preservation tokens carry roughly equal weight.

The Animate Anyone paper (Hu et al., 2023, arxiv.org/abs/2311.17117) established the foundational architecture for character-locked motion synthesis. Identity features and motion features benefit from being treated as separate signal streams with explicit preservation constraints.


FAQ

What is the difference between character lock in I2V and Ref2V? In I2V, the model sees one still and continues it without re-imagining the subject. In Ref2V, the model sees references and composes a new first frame. I2V locks the entire starting state pixel-for-pixel; Ref2V locks identity only, freeing lighting, pose, and composition.

Does character lock work better with a real photo or a generated image? A real photo under neutral lighting locks better than a generated image at the same resolution. Generated images often contain subtle identity-distorting artifacts (face asymmetry, inconsistent ear shape, melted earring geometry) the model treats as legitimate anchors. QA generated stills before use.

What is the cheapest way to improve character lock without changing models? Improve the starting image. A sharper, front-facing, single-subject, neutrally posed, high-contrast source image produces stronger locks than any prompt tweak.

Can character lock survive a 180-degree camera orbit? Weakly. Most models hold identity through about 90 degrees of arc, then drift in the back half because the rear of the head provides less signal. For 360-degree shots, hard-lock on visible accessories (hair clip, jacket logo, watch). Kling 3.0’s elements panel handles this best.

How do I keep wardrobe stable across multiple clips? Register the wardrobe as a separate anchor — a clear outfit still or a Kling 3.0 wardrobe element. Reference it with "preserve wardrobe exactly" and "no wardrobe swap". If the outfit has a distinctive accessory, name it in the prompt for a verbal anchor.

What is the relationship between lock strength and motion quality? They trade off. Hard locks flatten motion into stiffness; soft locks prioritize motion realism at the cost of drift. Brand safety wants hard locks; music video B-roll often wants soft. Test both ends on your specific clip.

Are there character lock limits specific to Seedance 2.5, Veo 3.1, Kling 3.0, or Wan 3.0? Seedance 2.5 is strongest for multi-reference casts (see Seedance multi-reference image prompts). Veo 3.1 is strongest for single-subject narrative with audio. Kling 3.0 is strongest for face + wardrobe panel-registered subjects. Wan 3.0 is strongest for endpoint fidelity, weakest for mid-clip drift.


Conclusion

Character lock in I2V is not a feature toggle. It is a contract between the starting image, the lock-strength modifiers, and the model’s preservation behavior. The twelve templates above are starting points — every production has its own brand, casting, and drift tolerance. Run a 4-clip test grid (one per model, all else equal) to find each model’s lock ceiling for your subject.

For broader strategies, see the AI character consistency prompts in 2026 companion article. For Seedance tactics, see the Seedance 2.5 prompts guide and the Seedance multi-reference image prompts article.


Reviewed by the videosprompt.org editorial team · October 2026

Share Article

Related Articles

Recommended Reading

Ready to Get Started?

Experience our product immediately and explore more possibilities.