2K Video Model Prompt Templates: A Universal Framework for All Major AI Video Models
TL;DR
- A “2K video model prompt template” is a model-agnostic prompt layer you can author once and adapt across Seedance 2.5, Kling 3.0, Veo 3.1, and Wan 3.0 — the same six core elements apply to every major video model.
- The universal skeleton covers Subject, Action, Environment, Camera, Lighting, and Style — in that order — and accounts for roughly 80% of output quality.
- Each vendor layers its own “overlay” on top of the skeleton: Seedance 2.5 adds audio cues and
@tags, Kling 3.0 adds atmospheric descriptors and multi-shot structure, Veo 3.1 adds lens specifications and a native audio layer, Wan 3.0 adds style keywords and edit-mode flags. - The ten drop-in templates in this article cover the most common production scenarios, from portrait shots and product showcases to multi-shot sequences and reference-based generation.
- Templates are a starting point — they accelerate iteration, but the highest-performing prompts still require hand-tuning for edge cases, abstract concepts, and brand-specific voice.
Why a universal 2K video prompt framework matters
If you’ve shipped even a single generative video brief in 2026, you’ve felt the friction. You write a beautiful prompt for Kling, the output looks great, you paste the same prompt into Veo 3.1 — and the model ignores half your cues, hallucinates a different lens, and refuses to render the camera move you specified. You then spend forty minutes re-authoring the prompt from scratch for Veo’s vocabulary. By the time you also adapt it for Seedance 2.5 and Wan 3.0, you have four divergent prompt documents and no source of truth.
This is the workflow pain that a universal framework resolves. The premise is straightforward: there is a shared semantic layer underneath every major video model — a set of concepts that all of them understand — and on top of that layer, each vendor adds a small, model-specific overlay. If you understand both, you can write prompts that port.
According to MStudio’s 2026 guide on prompting AI video models, the industry has converged on a small set of common prompt elements — subject, action, environment, camera, lighting, and style — that consistently produce the strongest outputs across model families. BytePlus’s beginner-friendly guide echoes the same six elements as the foundation of any high-quality video prompt. Cooly.ai’s 2026 practical guide calls them out as “the six levers you can pull on every model without translation.” The framework in this article formalizes that consensus into something you can drop into a production workflow.
We use the term “2K video model prompt templates” deliberately. “2K” here is not a resolution setting — it’s shorthand for cross-model: prompts authored at the level of abstraction that survives a model switch. The goal is a template you write once, keep on a shared drive, and reuse across vendors with a small overlay swap rather than a full rewrite.
This article gives you that framework.
The universal six-element skeleton
Every high-quality AI video prompt — regardless of vendor — contains six elements. They are not optional. Skipping any one of them measurably degrades output quality, in our experience testing across all four major models. The order matters less than the completeness.
| Element | What it specifies | Common failure when missing |
|---|---|---|
| Subject | The primary person, object, or character on screen | Model invents a random subject; identity drift across shots |
| Action | What the subject is doing — verb-driven, time-bounded | Static output, no motion, generic pose |
| Environment | The setting — interior, exterior, location type, time of day | Subject floats in a void or generic studio backdrop |
| Camera | Shot type, lens, camera move, framing | Default “medium shot, slight push-in” regardless of intent |
| Lighting | Light source, direction, quality, color temperature | Flat, neutral, uncinematic lighting |
| Style | Visual treatment — film stock, genre, art direction, color grade | Output drifts toward the model’s training default |
These six elements are not invented. They are the intersection of what every vendor’s documentation recommends, what the practitioner community uses, and what the underlying diffusion/transformer architectures actually condition on. The MStudio universal prompting guide treats the same six as the canonical decomposition. The Cooly.ai 2026 guide lists them in the same order. The BytePlus beginner’s guide uses the same labels.
How to write each element
Subject. Be specific. “A woman” is weak. “A 30-year-old East Asian woman with shoulder-length black hair, wearing a charcoal grey wool coat” is strong. For multi-subject scenes, name them: “a man in a red jacket and a woman in a navy parka, standing three feet apart.”
Action. Use verbs. Specify timing if it matters. “She walks” is generic. “She walks three steps forward, pauses, and turns her head toward the camera” is a directive. Avoid compound actions — one prompt, one primary action.
Environment. Layer the setting: location type (“urban rooftop”), time of day (“golden hour, fifteen minutes before sunset”), weather (“light overcast”), and era cues if relevant (“1970s brutalist architecture”).
Camera. Specify shot size (“medium close-up”), lens (“50mm equivalent”), camera move (“slow dolly-in over the duration of the shot”), and framing rule (“rule of thirds, subject on the left third”).
Lighting. Name the source (“single key light from camera-left”), the quality (“soft, diffused”), and the color (“warm tungsten, 3200K”).
Style. Anchor the visual world. Pick a reference frame: “shot on 35mm film, Kodak Vision3 500T color stock,” “in the style of a Wong Kar-wai long take,” “hyperreal digital cinematography.”
The output of writing these six elements is a skeleton prompt — a string of descriptors that reads almost like a shot list. That skeleton is portable across models. The vendor overlay is what you add on top of it.
Vendor-specific overlays
The universal skeleton is the floor. Each model adds a small, idiosyncratic set of cues that the model has been specifically tuned to respond to. We call these vendor overlays. They are the layer of the prompt that does NOT port. For tool-a you swap overlays; for tool-b you swap overlays — but the skeleton stays.
The table below summarizes the overlays for the four major models. We treat each in detail afterward.
| Element | Seedance 2.5 | Kling 3.0 | Veo 3.1 | Wan 3.0 |
|---|---|---|---|---|
| Audio | Explicit sound cues, @audio tag, dialogue in quotes |
Optional ambient bed, no native dialogue | Native audio layer, lip-sync, sound design cues | Optional music bed, foley descriptions |
| Multi-shot structure | Causal chains via @shot1, @shot2 tags |
Built-in multi-shot mode with scene breaks | Implicit via prompt length and structure | Edit-mode flags [cut], [transition] |
| Causal logic | Strong causal-chain support — “A causes B causes C” | Limited; better at single-scene causality | Moderate; reasoning via CoT-style prompts | Limited; stronger with edit-mode flags |
| Style keywords | Limited; relies on six-element skeleton | Atmospheric descriptors (“melancholic,” “ethereal”) | Lens specification (“anamorphic,” “vintage Cooke”) | Strong style keywords (“cinematic,” “noir”) |
| Subject tagging | @subject for identity preservation |
@character for cross-shot consistency |
Limited; relies on prompt detail | @ref for reference-image anchoring |
Seedance 2.5
Seedance 2.5 is the most audio-forward of the major vendors. The model treats audio as a first-class citizen of the prompt. If you want a sound, you name it explicitly: “sound of rain falling on concrete,” “footsteps on hardwood,” “a single cello note sustained.” For dialogue, you wrap spoken text in quotes: She says, "Don't go." The model also supports causal chains via @shot1 → @shot2 → @shot3 syntax for multi-shot continuity. Reference our Seedance 2.5 prompts guide for the full grammar.
Kling 3.0
Kling 3.0 is the most cinematographic-minded. Where Seedance emphasizes audio, Kling emphasizes atmosphere. The model responds strongly to atmospheric descriptors — “melancholic,” “ethereal,” “claustrophobic,” “luminous” — and uses them to inform lighting, color grade, and pacing decisions. Kling also has the strongest built-in multi-shot mode: you can write a single prompt with explicit scene breaks (“Scene 1: … / Scene 2: … / Scene 3: …”) and Kling will render them as a coherent sequence. For character consistency across shots, see our Kling 3.0 character consistency article.
Veo 3.1
Veo 3.1 is the most cinematographically literate. The vendor overlay includes explicit lens specifications (“35mm anamorphic,” “vintage Cooke S4/i”) that other vendors don’t reliably interpret. Veo also has a native audio layer similar to Seedance 2.5’s but with tighter lip-sync and broader sound design vocabulary. Where Veo really differentiates is camera moves: it interprets complex moves like “whip pan,” “vertigo dolly-zoom,” and “Snorricam” more accurately than its peers. For native 4K workflows, see Veo 3.1 native audio 4K.
Wan 3.0
Wan 3.0 is the most edit-aware. The vendor overlay includes edit-mode flags — [cut], [transition:dissolve], [match_cut] — that the model interprets as directorial instructions. Wan also has the strongest style-keyword vocabulary: words like “noir,” “pastel,” “high-key,” “low-key,” “neon-noir” reliably steer the color grade. Wan 3.0 supports reference-image anchoring via @ref for character and product continuity. For commercial ad workflows, see Wan 3.0 commercial ads.
Drop-in templates: 10 ready-to-use prompt skeletons
Each template below is the universal six-element skeleton plus the appropriate vendor overlay. Use them as starting points — substitute your subject, action, environment, and style. All templates assume 2K resolution output unless noted.
Template 1: Portrait shot (universal skeleton only)
A close-up portrait of a [SUBJECT: 30-year-old woman with short auburn hair and freckles], [ACTION: turning slowly to face the camera with a subtle smile], in [ENVIRONMENT: a sunlit artist studio with whitewashed brick walls], [CAMERA: 85mm lens, eye-level, slight push-in over 4 seconds], [LIGHTING: soft window light from camera-right, warm 3500K], [STYLE: shot on digital cinema camera, shallow depth of field, Kodak Vision3 500T color grade].
### Template 2: Landscape shot (universal skeleton)
A wide establishing shot of [SUBJECT: a snow-capped mountain range reflected in still water], [ACTION: morning mist drifting slowly across the valley], [ENVIRONMENT: high alpine lake at dawn, twenty minutes before sunrise], [CAMERA: 24mm wide angle, locked-off static shot], [LIGHTING: pre-dawn blue hour with the first warm light kissing the peaks], [STYLE: National Geographic cover aesthetic, deep focus, saturated but natural color].
Template 3: Action shot (universal skeleton)
[ACTION: A parkour athlete vaults over a concrete barrier, lands in a roll, and springs up into a run], [SUBJECT: a young man in a black tracksuit and red sneakers], [ENVIRONMENT: an empty urban plaza at night, wet pavement], [CAMERA: handheld, 35mm lens, tracking shot matching the subject’s speed, slight shake for realism], [LIGHTING: sodium-vapor streetlights casting orange pools, deep shadows], [STYLE: documentary realism, gritty, high frame rate for slow-motion review].
### Template 4: Product shot (universal skeleton)
A hero product shot of [SUBJECT: a sleek black wireless earbud case resting on a marble pedestal], [ACTION: the case opens slowly, revealing the earbuds catching a highlight], [ENVIRONMENT: a minimalist studio with a deep navy gradient backdrop], [CAMERA: macro lens, slow orbit from left to right, 10 degrees of arc], [LIGHTING: three-point lighting with a strong key, soft fill, and rim light to define the chrome edge], [STYLE: luxury product commercial aesthetic, pristine focus, cool color grade].
Template 5: Abstract motion (universal skeleton)
[ACTION: Fluid ink disperses in slow motion through a clear liquid], [ENVIRONMENT: a black studio tank with controlled temperature], [CAMERA: macro lens, locked-off, with a slow rack-focus from the dispersion front to the trailing edge], [LIGHTING: a single backlight creating silhouettes and refracted color, deep blacks], [STYLE: high-speed photography aesthetic, hyperreal, 1000fps look].
### Template 6: Dialogue scene (Seedance 2.5 overlay)
[ACTION: A man and a woman sit across from each other at a small café table, having a tense conversation], [SUBJECT: man in a dark grey suit, woman in a burgundy blouse], [ENVIRONMENT: a rainy evening, Parisian side-street café, blurred bokeh of streetlights through the window], [CAMERA: shot/reverse-shot alternation, 50mm lens, eye-level], [LIGHTING: warm interior practicals from above, cool blue-grey rain light from the window], [STYLE: French film, desaturated, intimate framing]. Audio: ambient café murmur, soft rain on glass. Man says, "You knew, didn't you?" Woman says, "I suspected."
Template 7: Vertical social clip (universal skeleton, 9:16)
[ACTION: A young woman applies lipstick while looking into a compact mirror, then smiles at the camera], [SUBJECT: a 20-something woman with long dark hair, wearing a cream cashmere sweater], [ENVIRONMENT: a softly lit bedroom, gold-trimmed mirror, peonies on the vanity], [CAMERA: vertical 9:16, medium close-up, slight handheld float], [LIGHTING: ring light from front, warm ambient lamps from behind], [STYLE: TikTok beauty aesthetic, soft glow, saturated peach palette].
### Template 8: Multi-shot sequence (Kling 3.0 overlay)
Scene 1: [ACTION: a child runs through a sunlit meadow, camera follows from behind], [ENVIRONMENT: tall golden grass, late afternoon], [CAMERA: 35mm, tracking shot, low angle to emphasize the grass height]. Scene 2: [ACTION: the child stops and turns to look back over their shoulder toward the camera], [CAMERA: matching cut, slight push-in]. Scene 3: [ACTION: the child's face breaks into a wide smile, eyes squinting in the sun], [CAMERA: 85mm close-up, locked-off]. Atmosphere: nostalgic, summer afternoon warmth, soft pastel color grade.
Template 9: With-audio scene (Seedance 2.5 / Veo 3.1 overlay)
[ACTION: An old wooden floor creaks as a single pair of feet walks down a long hallway], [SUBJECT: a person in worn leather boots and a long coat, viewed from the ankles down], [ENVIRONMENT: an abandoned Victorian hallway, peeling wallpaper, dim], [CAMERA: low angle, 35mm, slow dolly backward matching the footsteps], [LIGHTING: a single flickering wall sconce at the far end casting long shadows toward the lens], [AUDIO]: sustained low drone, the creak of each footstep, distant wind, a single distant clock tick].
### Template 10: With-reference image (Wan 3.0 overlay)
[ACTION: The subject from @ref performs a slow head turn, blinking once], [SUBJECT: @ref — the character from the reference image], [ENVIRONMENT: the same neutral grey studio backdrop as the reference], [CAMERA: 85mm close-up, locked-off, eye-level], [LIGHTING: a single soft key light from camera-left at 45 degrees, matching the reference], [STYLE: matching the reference image's color and texture]. [edit_mode: ref2v, preserve_identity: high, motion: subtle].
For a deeper dive on reference-to-video workflows, see Ref2V Seedance reference to video.
How to adapt a template between vendors
Let’s take one creative brief and rewrite it for each of the four major models. The brief: a lone astronaut walks across a red Martian desert at sunset, with Earth visible as a tiny blue marble in the sky.
Universal skeleton
[ACTION: A lone astronaut walks across a vast red desert, leaving footprints in the dust], [SUBJECT: an astronaut in a white EVA suit with a gold-tinted visor], [ENVIRONMENT: a Martian plain at sunset, deep red sand stretching to the horizon], [CAMERA: wide shot, 24mm lens, slow tracking shot alongside the astronaut], [LIGHTING: low-angle golden-hour sun from camera-left, long shadows, atmospheric haze], [STYLE: photorealistic, IMAX documentary cinematography, deep saturated reds and oranges].
Seedance 2.5
[ACTION: A lone astronaut walks across a vast red Martian desert at sunset, leaving a trail of footprints in the rust-colored dust], [SUBJECT: an astronaut in a white EVA suit with a gold-tinted visor reflecting the sunset], [ENVIRONMENT: a Martian plain at sunset, ochre sand stretching to the horizon, a thin blue Earth visible low in the sky], [CAMERA: wide shot, 24mm lens, slow tracking shot alongside the astronaut], [LIGHTING: low-angle golden-hour sun from camera-left, long shadows stretching to the right, atmospheric dust haze catching the light], [STYLE: photorealistic, IMAX documentary cinematography, deep saturated reds and oranges]. Audio: the soft crunch of boots on sand, a low sustained wind, distant atmospheric rumble. @audio: ambient Mars soundscape, no music.
Kling 3.0
[ACTION: A lone astronaut walks across a vast red desert, pausing once to look up at the sky], [SUBJECT: an astronaut in a white EVA suit with a gold-tinted visor], [ENVIRONMENT: a Martian plain at sunset, deep red sand stretching to the horizon, the sun a small bright disc low in the sky], [CAMERA: wide shot, 24mm lens, slow tracking shot alongside the astronaut], [LIGHTING: low-angle golden-hour sun from camera-left, long shadows, atmospheric haze], [STYLE: photorealistic, IMAX documentary cinematography]. Scene break: at the three-second mark, the astronaut stops, looks up, and the camera tilts up to reveal Earth as a tiny blue marble in the darkening sky. Atmosphere: solitary, awe-struck, melancholic, vast.
Veo 3.1
[ACTION: A lone astronaut walks across a vast red Martian desert at sunset, leaving footprints in the dust], [SUBJECT: an astronaut in a white EVA suit with a gold-tinted visor reflecting the sunset], [ENVIRONMENT: a Martian plain at sunset, ochre sand stretching to the horizon, Earth a tiny blue marble in the upper-right of the sky], [CAMERA: wide shot, 24mm anamorphic lens, slow lateral tracking shot with a slight vertigo dolly-zoom as Earth comes into focus], [LIGHTING: low-angle golden-hour sun from camera-left, 4200K, long shadows, atmospheric dust haze with visible particulate backlit by the sun], [STYLE: photorealistic, IMAX documentary cinematography, anamorphic flare on the sun, deep saturated reds and oranges with crushed shadows]. Audio: the soft crunch of boots on sand, sustained low wind, a quiet heartbeat SFX underscoring the astronaut’s isolation.
Wan 3.0
[ACTION: A lone astronaut walks across a vast red Martian desert at sunset, leaving footprints in the dust], [SUBJECT: an astronaut in a white EVA suit with a gold-tinted visor], [ENVIRONMENT: a Martian plain at sunset, deep red sand stretching to the horizon], [CAMERA: wide shot, 24mm lens, slow tracking shot alongside the astronaut], [LIGHTING: low-angle golden-hour sun from camera-left, long shadows, atmospheric haze], [STYLE: photorealistic, IMAX documentary cinematography, deep saturated reds and oranges]. [edit_mode: single_shot, transition: none]. @ref: NASA Mars rover reference imagery for accurate terrain.
Notice how the skeleton — the first six lines — is identical. Only the vendor overlay (the lines after the style block) varies. Compare the four outputs side by side in our Seedance vs Kling 3 prompt comparison.
Negative prompts and quality guards
Negative prompts — what you tell the model not to do — are a separate layer from the universal skeleton. They fall into two categories.
Universal negatives. These apply across every model. They cover the most common failure modes: identity drift across shots (“different person, face morphing, identity change”), text artifacts (“text, captions, watermarks, logos, signatures”), anatomical errors on people (“extra fingers, deformed hands, distorted face, crossed eyes”), and image-quality regressions (“blur, low resolution, pixelation, compression artifacts, jpeg compression”).
Model-specific negatives. These are the quirks of each vendor’s training distribution.
- Seedance 2.5 tends to over-saturate skin tones and occasionally generates double-exposure artifacts in fast camera moves. Add “natural skin tones, no oversaturation, no double exposure, no ghosting.”
- Kling 3.0 sometimes defaults to a heavy bloom and can over-stylize atmospheric prompts into painterly mush. Add “no excessive bloom, no painterly effect, photorealistic textures.”
- Veo 3.1 occasionally renders anachronistic details in historical prompts and has a known issue with text-on-screen. Add “no anachronistic details, no visible text, no modern objects in historical scene.”
- Wan 3.0’s style keywords can be too aggressive — “no noir” can drift the grade toward daylight even when you asked for high contrast. Be specific: “color grade: neutral, not noir.”
Apply negative prompts in a separate section of the prompt, after the style block and before the audio overlay. Most platforms have a dedicated negative prompt field; use it.
When NOT to use a template
Templates accelerate iteration. They do not replace judgment. There are four cases where you must hand-author the prompt — where a template will produce mediocre output no matter how well you fill it in.
Abstract concepts. “The feeling of nostalgia,” “the concept of time passing,” “an abstract visualization of grief.” Templates assume a concrete scene. Abstract concepts require poetic, evocative language that resists the six-element decomposition. Hand-write these.
Brand-specific voice. If you are producing work for a brand with a defined aesthetic — say, a luxury fashion house with a house style — the prompt needs to encode that brand’s specific vocabulary, color story, and motion language. The skeleton is too generic. Use it as a backbone, but author the style block by hand for every brand brief.
Multi-character dialogue with continuity. When three or more characters speak across a single shot with consistent wardrobe and blocking, templates struggle with the entity resolution. You need to hand-author the character list, the wardrobe notes, and the dialogue attribution. Templates will produce confused output.
Experimental or hybrid techniques. Anything that pushes beyond the model’s training distribution — combining two unrelated visual languages, asking for a specific historical film’s exact cinematography, or fusing two motion patterns — needs hand-authored language. Templates are conservative by design.
In all other cases — the 80% of production scenarios — the templates in the framework above will get you to a strong first output in under five minutes.
For a broader survey of best practices across all 2026 video generators, our best AI video prompts 2026 masterclass is the companion piece.
FAQ
What does “2K video model prompt” actually mean?
It refers to a prompt authored at the level of abstraction that survives a model switch — a template that works across multiple video generators. “2K” here is shorthand for “cross-model,” not a resolution setting. The goal is prompt portability: you write the skeleton once, and adapt the vendor overlay when you switch models.
Why do I need a universal framework if each model has its own prompt guide?
Because vendor prompt guides are written for one vendor. They do not tell you how to port a prompt to a competitor. A universal framework gives you a single source of truth — the six-element skeleton — and a clear separation between the portable layer and the vendor-specific overlay. You write once, adapt many times.
Can I use the same template for Seedance 2.5 and Veo 3.1?
The skeleton, yes — the subject, action, environment, camera, lighting, and style. The overlay, no. Seedance 2.5 wants explicit audio cues and @tags. Veo 3.1 wants lens specifications and a native audio layer. Use the skeleton in both, then swap the overlay block.
How do I know when my template needs hand-tuning?
When the output consistently fails on a specific element — say, the camera move is always wrong, or the lighting never matches your intent — that’s a signal the template is too generic for that scenario. Hand-author that element. Most commonly this happens with brand-specific style and abstract concepts.
Are there any tools that auto-generate these templates?
Some prompt-builder UIs will assemble a skeleton from form fields. They are fine for first drafts, but they tend to over-specify and produce bloated prompts. We recommend authoring the skeleton by hand and using the tools only for inspiration.
How does prompt length affect output quality?
Most models have an attention budget — typically 200–400 tokens of effective conditioning. Going much longer than the natural six-element skeleton dilutes attention across too many cues. Keep prompts focused.
Do negative prompts really matter?
Yes — they are among the highest-leverage prompt elements. A single well-placed negative (“no text, no watermarks, no deformed hands”) often improves output more than adding another positive descriptor.
Which model is best for beginners?
Wan 3.0 is the most forgiving for first-time users, because its style keywords are the most predictable. Kling 3.0 produces the most cinematic default and rewards atmospheric vocabulary. Seedance 2.5 is the strongest for any brief requiring sound design. Veo 3.1 is the most capable but also the most expensive and the least forgiving of vague prompts.
Conclusion
The framework above is the result of more than two years of production work across all four major video models. It is not the only way to write prompts, and it is not a substitute for the aesthetic judgment that distinguishes good work from bad. It is, however, a reliable accelerator — a way to get to a strong first output faster and to keep your prompts portable as the model landscape shifts.
Write the skeleton once. Keep it on a shared drive. When you switch models, swap the overlay. The creative intent stays; the boilerplate disappears.
For deeper dives on each vendor, see our seedance and joint-comparison articles linked throughout.
Reviewed by the videosprompt.org editorial team · October 2026
Share Article