VideosPrompt VideosPrompt

How to Convert a Reference Video into an AI Video Prompt: A Workflow for Seedance, Veo, Kling, and Wan

Author: VideosPrompt Date: 2026-10-07 14:11:44
How to Convert a Reference Video into an AI Video Prompt: A Workflow for Seedance, Veo, Kling, and Wan

TL;DR

  • Every AI video prompt is built from four extractable layers: shot type and camera movement, subject and action, environment and lighting, and style plus post-process feel.
  • Frame-by-frame annotation in a free tool (VLC, DaVinci Resolve, or OBS) is the fastest way to capture those four layers as discrete observations before you write a single word of prompt.
  • Camera-motion extraction is the highest-leverage step — describing dolly, orbit, pan, push, and crane moves with consistent verbs produces dramatically more reproducible output than vague “slow cinematic” language.
  • A reference video becomes a working prompt once you collapse your notes into a single dense paragraph that pairs each camera verb with a subject, an environment, and a grade; we walk through a perfume-ad worked example below.
  • Prompt diffing — comparing your generated output to the reference shot by shot — is the loop that turns a one-off generation into a reusable template; the 12 templates at the bottom of this article are pre-diffed.
  • This workflow is model-agnostic; the same extraction process feeds Seedance, Veo 3.1, Kling 2.x, and Wan 2.x with only minor vocabulary adjustments.

The reverse-engineering problem

Every filmmaker has the same moment: you see a clip and want to make something that captures its energy — not a copy, but the same camera, the same mood. In 2026, the fastest path from inspiration to output runs through a text-to-video model. The hard part is not the model; it is the prompt.

Reverse-engineering videos is hard because language and cinematography speak different grammars. A camera move is one thing to a cinematographer and another to a prompt parser. A “soft afternoon” is one thing to a director of photography and another to a diffusion model. The workflow below treats the reference not as inspiration but as a spec sheet — a document you can read, decompose, and rewrite in the native vocabulary of whichever model you are driving.

The use cases are wider than they look:

  • Studying references. Junior creators can build a personal library of prompts decoded from masters, learning cinematography through prompt composition.
  • Recreating trends. When a visual style goes viral, prompt-level reverse-engineering lets you ship matching content in hours.
  • Building prompt libraries. Agencies can maintain shared prompt sets so a look developed once by one artist is reused across campaigns by another.
  • Reproducibility. A prompt derived from extraction is reproducible. A prompt inspired by vibes is not.

This article is a working procedure. Read it once end-to-end, then keep it open while you run the workflow on your first reference.

The four extractions

A video is a stack of decisions. To reverse-engineer one, separate the stack into four layers and write each one down independently. Every prompt contains all four.

1. Shot type and camera movement. What the camera is doing over time — wide establishing, medium close-up, macro insert, push-in, pull-back, orbit, dolly-track, handheld drift, static lock-off, aerial descent. Camera movement is the layer most often lost in amateur prompt writing, and it is the layer that distinguishes a real cinematic look from a slideshow of generative stills.

2. Subject and action. Who or what is in the frame, and what are they doing? A subject is “a woman in a tailored beige coat.” An action is “walking toward camera across an empty plaza.” The combination — subject doing action, in present tense — gives the model something to anchor motion around. The most common mistake here is too much subject detail (race, age, body type, clothing brand) without enough action detail.

3. Environment and lighting. Where are we, when, and what is the light doing? Indoor vs. outdoor, urban vs. natural, day vs. night, hard sun vs. overcast softbox, golden hour vs. blue hour. Lighting is the cheapest mood lever in any video, but it is also described most often with adjectives (“beautiful,” “moody”) rather than concrete cues (“key light from camera-left at 45 degrees, warm 3200K, deep shadow on the right side of face”).

4. Style and post-process feel. What is the final image doing beyond the captured scene? Color grade (teal-and-orange, bleach bypass, low-contrast pastel), lens character (anamorphic flare, shallow DOF, vintage spherical softness), film emulation, grain structure, frame rate (24 fps cinematic vs. 30 fps broadcast vs. 60 fps hyperreal). This is where most “look” decisions actually live.

Capture all four, even when one feels obvious. A static close-up in golden hour still has a shot type, a subject, an environment, a style — and writing each one down keeps you honest about what is in the reference.

Frame-by-frame annotation

Annotation is the bottleneck, and the step most people skip. Do not skip it. A thirty-second reference contains roughly 720 frames at 24 fps; trying to absorb that much visual data in one pass is impossible, and any prompt you write from a single impression will be vague.

Pick a free playback tool that lets you scrub frame-by-frame and write notes alongside:

Tool Best for Key feature for annotation
VLC Quick single-pass review Frame-by-frame stepping with E key, frame export as PNG
DaVinci Resolve Reference library for ongoing projects Built-in metadata fields per clip, marker system for in/out points
OBS + media source Recording your own reference library Capture reference clips with chapter markers for later scrubbing

For a 30-second reference, plan on 15-30 minutes of annotation in passes:

  1. Rhythm. Watch the whole clip at speed. Note the cuts, transitions, and overall arc. How many distinct shots? Where does energy shift?
  2. Per-shot breakdown. Scrub shot by shot. For each, write the four layers in shorthand: “Shot 3 — 00:07 to 00:11 — MCU, slow push-in. Woman in beige coat, walking. Outdoor plaza, late afternoon, warm side-light from screen-right. Style: shallow DOF, slight bloom, muted grade.”
  3. Outliers. Return to any shot that felt different — usually the opening and closing shots — and re-tag them with the same care.

The output is a list of cards, one per shot. That list is your prompt source-of-truth. Everything that follows is mechanical: compressing a card into a sentence, comparing output to a card, folding several cards into a single multi-shot prompt.

Camera-motion extraction

Camera movement is the highest-leverage layer, because it is the layer that most generative models actually understand well. Describe the camera precisely and the model will compose accordingly; leave it vague and the model will guess, rarely matching the reference.

The vocabulary you use matters. Different prompt parsers weight different words differently, but the verbs below are widely recognized across Seedance, Veo 3.1, Kling 2.x, and Wan 2.x.

Verb What it means When to use it
Static / locked-off Camera does not move Product shots, portraits, contemplative moments
Pan Camera rotates horizontally on a fixed axis Reveals, follows horizontal motion
Tilt Camera rotates vertically on a fixed axis Reveals height, follows vertical motion
Dolly Camera physically moves forward or backward Push-ins for intimacy, pull-backs for reveal
Truck Camera physically moves left or right Parallax, lateral follows
Pedestal Camera physically moves up or down Vertical reveals, height transitions
Orbit / arc Camera moves in a circular path around subject Hero product shots, dramatic reveals
Crane / boom Camera moves through space on a sweeping arc Epic reveals, transitions, aerials
Handheld Camera carried by operator, visible micro-movement Documentary feel, urgency
Steadicam / gimbal Smooth moving camera, locked horizon Walking shots, fluid follows
Aerial / drone High vantage, often with slow movement Establishing shots, landscape

Three parameters sharpen the description:

  • Speed. Slow, medium, fast, or a duration (a “two-second push-in” is more useful than “slow push-in”).
  • Direction. Toward subject, away from subject, screen-left to screen-right, ascending, descending.
  • Frame impact. Does the camera start wide and end tight? Stay locked? The destination framing often matters more than the path.

Write the camera move first, before the subject. “Slow push-in on” then a comma then the subject reads more like a shot list and is parsed more reliably than the other way around. The general prompt-engineering principles outlined in mstudio.ai’s guide to prompting AI video models reinforce this ordering — the model latches onto the first concrete noun phrase it sees.

From extraction to prompt: a worked example

To make the workflow concrete, walk through a 30-second reference — a hypothetical perfume ad. You have already done the annotation passes above; you have a stack of six shot cards.

Shot Timecode Camera Subject + action Environment + light Style
1 00:00-00:04 Slow push-in, medium close-up Perfume bottle rotating slowly on a black-glass plinth Studio, single key light from above, deep shadow High contrast, anamorphic streak on highlight
2 00:04-00:09 Orbit, 180 degrees Same bottle, full-body in frame, slow rotation Same studio, rim light becomes visible from behind Same
3 00:09-00:14 Cut to extreme close-up, slow tilt up Glass facets catching light, micro-droplets on the surface Same Same, with macro bokeh
4 00:14-00:19 Cut to medium, static A woman’s hand reaches in from screen-right, picks up bottle Same Same
5 00:19-00:25 Crane up and back, starting at the hand, ending wide on the plinth and surrounding dark studio Hand lifts bottle out of frame at the top of the move Same, but the wide reveals a single overhead spotlight beam Same, with slight bloom
6 00:25-00:30 Static wide, slow zoom out Black frame with logo fades up over the plinth Same Same

To turn those six cards into a single Seedance prompt, do not write all six as one prompt. Generative video models do not reliably chain shots; multi-shot narratives remain a weakness across the field. Instead, choose the one shot that captures the most of the reference’s intent and write that as your prompt, then generate variations.

Shot 2 — the orbit — is the right choice. It carries the hero subject (the bottle), the signature camera move (orbit), the lighting signature (rim light), and the style cue (anamorphic) in one beat. Here is the prompt derived from that card:

Template — Perfume Bottle Orbit (Seedance 2.5 / Veo 3.1 / Kling 2.x / Wan 2.x)

Slow 180-degree orbit around a faceted glass perfume bottle on a black-glass plinth, full-body in frame, gentle continuous rotation. Single key light from above casting deep shadow below, warm rim light wrapping the bottle from behind, dark studio backdrop. Anamorphic lens character with horizontal streak on highlights, shallow depth of field, subtle bloom. Cinematic, high contrast, photoreal, 24fps, 4K.

Notice what is in the prompt and what is not. The verb (“orbit”) is first. The subject (“faceted glass perfume bottle”) is concrete and physical — not “luxury fragrance.” The lighting is described by position and color temperature, not adjectives. The style cues (anamorphic, shallow DOF, bloom) are listed at the end where they read as modifiers to everything before them. Nothing in the prompt is fabricated — every word traces back to a card.

For a workflow comparison and additional prompt-parsing conventions across models, Pixmind’s video-to-prompt workflow article is a useful parallel reference; the techniques translate even where the model differs.

Prompt diffing

The first generation from any extracted prompt will not match the reference. That is not a failure — it is the start of the workflow.

Prompt diffing compares your generated output to the reference shot by shot and asks, for each shot, three questions:

  1. What is the same? Identify elements that already match — camera move, lighting direction, grade. Do not change those.
  2. What is different but acceptable? Identify elements that drifted but are still consistent with intent. Note them; do not over-correct.
  3. What is wrong? Identify elements that need to change. These are usually subject-related (wrong shape, texture, scale) or environment-related (background drifted). Fix these next iteration.

Keep a diff log. A spreadsheet with columns for shot, current prompt language, observed output, revised prompt language lets you accumulate corrections across generations. After three or four rounds you will have a prompt that reliably reproduces the reference’s energy — and one you understand well enough to vary on purpose.

A common trap is to keep adding modifiers to force closer matches. If a prompt is not landing, the answer is rarely a tenth adjective. The answer is usually structural: rewriting the camera verb, repositioning the subject, or changing sentence order. When in doubt, return to your annotation cards and ask which layer your revision is targeting.

12 prompt templates

The templates below are pre-diffed for their respective styles. Each was derived from a real reference using the workflow above and refined through three to five generations. Use them as starting points; replace bracketed placeholders with your own subject matter.

Cinematic

Template 01 — Cinematic Establishing Wide

Locked-off wide shot of [LOCATION] at [TIME OF DAY], [SUBJECT] small in frame. [KEY LIGHT DESCRIPTION], deep shadow on screen-right. Anamorphic widescreen 2.39:1, desaturated grade, gentle film grain, 24fps.
Template 02 — Cinematic Push-In Portrait

Slow dolly push-in on [SUBJECT DESCRIPTION] in [LOCATION]. Medium close-up tightening to close-up over [DURATION]. Soft key light from camera-left at 45 degrees, fill at quarter stop, gentle background separation. Shallow depth of field, naturalistic color, 24fps.
Template 03 — Cinematic Handheld Follow

Handheld follow shot trailing [SUBJECT] from behind as they walk through [LOCATION]. Visible micro-movement, motivated natural light, ambient environmental sound implied. Documentary texture, slightly lifted blacks, 24fps.

Social

Template 04 — Vertical Social Reveal

Vertical 9:16 frame, slow pedestal move upward revealing [SUBJECT]. [LIGHT DESCRIPTION], saturated color, sharp foreground with creamy bokeh background. Modern, energetic, loop-friendly beat, 30fps.
Template 05 — Quick-Cut Trend Beat

Fast push-in to extreme close-up of [SUBJECT DETAIL], 1-second beat, hard cut to medium wide of full [SCENE]. Bold contrast, slightly overexposed highlights, saturated grade, social-native framing, 30fps.
Template 06 — Talking-Head Static

Locked-off medium close-up of [PRESENTER], direct eye contact with camera. Soft diffused key light, low-contrast background, minimal grade. Clean, trustworthy, broadcast-friendly, 30fps.

Product

Template 07 — Hero Product Orbit

Slow 180-degree orbit around [PRODUCT] on [SURFACE]. Full product in frame, gentle continuous rotation. Studio lighting with overhead key and rear rim, dark or gradient backdrop. Anamorphic lens character, horizontal highlight streak, shallow depth of field, photoreal, 24fps.
Template 08 — Macro Detail Insert

Extreme close-up of [PRODUCT DETAIL], slow tilt upward. Macro lens character, creamy bokeh, light catching surface texture. Studio lighting, high contrast, photoreal, 24fps.
Template 09 — Product In Context

Static medium shot of [PRODUCT] placed within [LIFESTYLE ENVIRONMENT]. Natural motivated light, props suggesting use, shallow depth of field isolating the product. Warm naturalistic grade, editorial composition, 24fps.

Narrative

Template 10 — Two-Shot Conversation

Static medium two-shot of [CHARACTER A] and [CHARACTER B] in [LOCATION]. Soft motivated practical lighting, shallow depth of field, naturalistic grade. Film dialogue coverage, 24fps.
Template 11 — Emotional Reveal

Slow crane-up from close-up of [SUBJECT'S HANDS OR OBJECT] to wide reveal of [LOCATION]. Continuous single take, motivated natural light, slightly muted grade, gentle bloom on highlights. Contemplative pacing, 24fps.
Template 12 — Closing Wide Pull-Back

Slow dolly pull-back starting medium on [SUBJECT], ending wide on [ENVIRONMENT]. Single take, fading light, deep shadow building on subject. Cinematic 2.39:1, low-contrast grade, gentle grain, 24fps.

For Veo 3.1-specific cinematography vocabulary — including anamorphic, focal-length, and audio-cue conventions — Google AI Studio Veo prompt cookbook is a strong companion. Several templates above use phrasing patterns that align well with Veo’s prompt parser.

Tools for the conversion

The workflow is deliberately tool-light. None of the steps require paid software, and the entire pipeline runs on a laptop you already own.

Observation tools. VLC for one-off scrubbing and frame export. DaVinci Resolve (free tier) for ongoing reference library management with built-in metadata. OBS Studio for capturing your own reference clips with chapter markers. Your browser’s developer tools for inspecting streaming-reference playback properties (resolution, frame rate, codec).

Prompt builders. Claude and ChatGPT as drafting partners — paste your annotation cards and ask for “a single dense prompt that captures shots 2 through 4 as if they were one continuous shot.” The model will not match your reference, but it will give you a starting paragraph to compress and edit.

Reference libraries. Pinterest boards remain the best tool for collecting still-frame references. Shotkit and Artgrid for curated cinematography references. For the AI-prompt layer, the Seedance 2.5 model guide and the Veo 3.1 native audio reference cover two of the strongest current models in detail.

Storage. A folder structure — references/raw, references/cards, references/prompts, references/diff-logs — pays for itself the third time you run the workflow. The most common mistake is keeping only the final prompt; the cards and diff logs are where the learning lives.

FAQ

Do I need to use the same model the reference video was generated with? No. The extraction workflow is model-agnostic. A prompt derived from a Seedance reference can be rewritten for Veo 3.1, Kling, or Wan with only minor vocabulary adjustments. The four layers are universal.

How long should the reference video be? Shorter is better for learning the workflow. Start with 15-to-30-second references before attempting longer pieces. Multi-shot, 60-second-plus references benefit from a more aggressive pre-selection of which one or two shots to model on.

What if I cannot describe the camera move precisely? Watch the reference with sound off, slowed to quarter speed. The camera move is usually the most legible element once you slow it down. If you still cannot name it, default to “slow dolly push-in” or “slow orbit” — both are widely recognized across models.

How do I handle references that mix multiple styles? Pick one style cue per generation. A reference that blends an anamorphic look with a vintage film grade will often produce muddy results if you try to land both at once. Generate with one, then re-generate with the other and pick the best frame from each as your final.

Is this workflow useful for image-to-video as well? Yes, with one adjustment: when feeding an image plus a prompt to an image-to-video model, the four-extraction workflow produces the prompt while the reference image provides the visual anchor. The prompt carries the motion, lighting, and style; the image carries the subject and composition. See our multimodal reference video prompts guide for details.

How many rounds of diffing does a good prompt usually take? Three to five is typical for a complex reference. If you are past ten rounds without convergence, the issue is usually structural — return to your annotation cards and re-extract the dominant layer rather than tweaking adjectives.

Are there universal templates I can reuse across styles? Yes — the 2K video model prompt templates collection covers the canonical 4K/2K framing patterns that work across Seedance, Veo, Kling, and Wan. Use the 12 templates in this article as style-specific starting points, and the 2K collection as your model-format reference.

Conclusion

Reverse-engineering a reference video into an AI prompt is not a single trick — it is a four-step discipline. Extract the four layers. Annotate frame by frame. Compress each card into a single sentence that leads with the camera move. Diff the output against the reference and revise with intent.

Run the workflow ten times and you will watch videos differently. The camera move becomes legible before the story does. The grade becomes a choice rather than a default. That is the real payoff — better visual literacy, which compounds across every model and every reference.

For broader foundations, the best AI video prompts 2026 masterclass is the right starting point; for model-specific deep dives, see the Seedance and Veo guides linked above.


Reviewed by the videosprompt.org editorial team · October 2026

Share Article

Related Articles

Recommended Reading

Ready to Get Started?

Experience our product immediately and explore more possibilities.