VideosPrompt VideosPrompt

Veo 3.1 Native Audio & 4K Product Video Prompts: The Complete Guide

Author: VideosPrompt Date: 2026-10-07 14:11:38
Veo 3.1 Native Audio & 4K Product Video Prompts: The Complete Guide

TL;DR

  • Veo 3.1 is Google DeepMind’s flagship video model that generates native audio (dialogue, SFX, ambience) and 4K resolution output in a single generation pass — no post-dubbing required (deepmind.google/models/veo).
  • The most reliable Veo 3.1 prompt formula is Subject → Action → Scene → Camera → Lens → Style → Audio → Negative, applied in that exact order.
  • Veo 3.1 responds strongly to cinematographer-grade vocabulary — “low angle,” “rack focus,” “tracking left,” “35mm lens,” and “Rembrandt lighting” all materially change the output (aistudio.google.com/docs/veo/cookbook).
  • Native audio has three distinct layers (dialogue, SFX, ambience). Each must be specified separately or Veo will improvise inconsistently.
  • 4K (3840×2160) output requires explicit selection in Vertex AI or the Gemini API; not all surface UIs expose it by default (Google AI Studio Veo prompt cookbook).
  • Veo 3.1’s known weak spots are duration control (8-second default clips) and multi-character synchronized dialogue — both addressed in the troubleshooting section below.

What Veo 3.1 is, and why native audio + 4K matters

Veo 3.1 is the third-generation video foundation model released by Google DeepMind, rolling out to Gemini API, Vertex AI, and Flow between October and November 2025. Unlike its predecessors, Veo 3.1 produces synchronized audio and 4K-resolution video from a single text prompt. The model is documented at the official Veo model page and the practical Veo prompt cookbook.

For product marketers, three capabilities change the workflow:

  1. Native audio generation. Dialogue, sound effects, and ambient sound are produced during the same inference pass as the video. You no longer need to lip-sync in Runway, ElevenLabs, or a DAW.
  2. 4K output. Veo 3.1 supports 4K (3840×2160) generation at multiple aspect ratios, suitable for hero web banners, broadcast cuts, and high-end e-commerce detail pages.
  3. Cinematographer-aware prompting. The model is trained on film terminology and responds meaningfully to lens, lighting, and camera-move vocabulary.

A November 2025 model comparison: Seedance vs Kling vs Veo positions Veo 3.1 against OpenAI’s Sora 2, noting Veo’s stronger audio fidelity and Sora 2’s longer default clip length. For product work, Veo’s audio layer is the deciding factor.

This guide is a working reference. It documents the prompt formula, the cinematographer’s vocabulary, the three-layer audio model, 15 production-ready templates across five product categories, and the failure modes that consistently waste credits.


The Veo 3.1 prompt formula

The most reliable structure for product video prompts is an eight-token chain. Order matters because Veo weights tokens by position. Earlier tokens dominate the composition; later tokens refine it.

Subject → Action → Scene → Camera → Lens → Style → Audio → Negative
Slot Purpose Example
Subject The hero product or model “A matte-black ceramic skincare bottle”
Action What happens in the clip “slowly rotates 90 degrees”
Scene The environment “on a wet slate slab in a forest”
Camera Movement and framing “slow dolly-in, eye level”
Lens Focal length and aperture behavior “85mm, shallow depth of field”
Style Color grade and look “warm tungsten, film-grain 35mm print”
Audio Dialogue, SFX, ambience “soft rain ambience, no dialogue, single droplet SFX at 2s”
Negative What to avoid “no text, no logos, no extra fingers, no harsh shadows”

The Vertex AI e-commerce prompt library (Google AI Studio Veo prompt cookbook) uses the same structure, though it condenses lens and style into a single “look” field. Splitting them gives finer control.


Speaking the cinematographer’s language

Veo 3.1 was trained on a corpus that includes cinematography references, so production-grade terms land. Generic phrasing (“a nice camera angle”) degrades the output. Specific phrasing materially improves it.

Camera moves that consistently work

  • Dolly-in / dolly-out — moving the whole camera toward or away from the subject. Reliable for product reveals.
  • Tracking left / tracking right — lateral movement parallel to the subject. Excellent for walking shots and conveyor reveals.
  • Crane up / crane down — vertical movement on a vertical axis. Useful for scale shots.
  • Low angle / high angle — point-of-view specification. Low angle makes products feel monumental; high angle feels editorial.
  • Rack focus — shifting focal plane between foreground and background. Veo 3.1 handles this for static compositions but struggles during fast movement.
  • Handheld — adds controlled shake. Use sparingly for lifestyle content; can read as “amateur” if the prompt also requests “cinematic.”
  • Orbit / arc — circular motion around the subject. The single most reliable move for product hero shots.

Framing terms

  • Close-up (CU), Medium Close-Up (MCU), Medium (MS), Wide (WS) — standard shot sizes.
  • Over-the-shoulder (OTS) — for dialogue or product demonstration with a presenter.
  • Insert / macro — extreme close-up for texture, ingredients, or detail storytelling.
  • Two-shot — for products shown with a human or paired object.

The Veo prompt cookbook confirms that combining a move with a framing term (e.g., “slow orbit, medium close-up”) yields more consistent results than either alone.


Native audio: dialogue, SFX, ambience

Veo 3.1’s audio model has three independent layers. Specifying each in your prompt — rather than writing “sound effects” generically — is the difference between a usable render and a clip with mismatched or missing sound.

Layer 1: Dialogue

Format the line in quotes and place it explicitly:

Audio: A female voiceover says, "Meet the new Hydra-Serum. Three drops, every morning."

For synced on-screen speech, name the speaker and give the lip-move cue:

Audio: A male presenter, off-camera, says, "Notice the brushed-aluminum finish."

Veo 3.1 handles single-speaker dialogue reliably. Two-speaker synchronized dialogue is one of its known weak spots (see below).

Layer 2: Sound effects (SFX)

Be specific. “Splash” produces a generic splash. “Single water droplet hitting marble at 1 second” produces a precise, cinematic hit.

Common useful SFX phrasings:

  • “Soft mechanical click of a magnetic closure”
  • “Single espresso machine pull, foam hiss at 2s”
  • “Subtle paper-tear sound”
  • “Glass placed gently on a coaster”

Layer 3: Ambience

Ambience is the bed — the room tone, the wind, the distant traffic. Without it, clips feel sterile.

  • “Quiet forest ambience, distant birdsong, no wind”
  • “Open-air rooftop terrace at golden hour, faint city hum”
  • “Soundstage interior, dead acoustic, no reverb”

Google AI Studio Veo prompt cookbook notes that specifying ambience explicitly yields noticeably more “produced” audio than relying on Veo’s default behavior.


Lens and lighting vocabulary that works

Lighting setups

Term Effect When to use
Rembrandt lighting Triangle of light under the eye, classic portrait Skincare, fragrance, hero portraits
Split lighting Half face lit, half in shadow Mystery, premium, masculine grooming
Butterfly / Paramount Light from above, small shadow under nose Beauty, glamour, jewelry
Rim / back light Edge highlight around subject Separating product from background
Golden hour Warm, low sun, long shadows Outdoor lifestyle, fashion
Overcast / soft Diffused, shadowless Skincare macro, food
Neon / practical Color from in-frame light sources Tech, gaming, nightlife

Lens and aperture

  • 35mm — standard wide, slight environmental context. Versatile.
  • 50mm — “natural” focal length, minimal distortion. The default for editorial product.
  • 85mm — flattering compression, classic portrait. The default for hero beauty shots.
  • Macro — extreme close-up, shallow depth of field, reveals texture.
  • Anamorphic — oval bokeh, horizontal flares. Cinematic but can distort product geometry.

Always pair lens with aperture behavior:

Lens: 85mm, shallow depth of field, f/1.8 equivalent

Without aperture, Veo defaults inconsistently.


15 production-ready product video templates

Each template below is formatted as a single prompt. Copy verbatim, then substitute your product, brand colors, or specific SFX.

Cosmetics and beauty

Template 1 — Serum drop hero

Subject: A 30ml matte-black glass serum bottle with a gold dropper cap.
Action: A single amber drop falls from the dropper onto the back of a hand.
Scene: On a white marble slab, soft window light from camera left, dust motes visible.
Camera: Slow dolly-in, macro framing, eye level.
Lens: 85mm, shallow depth of field, f/2.0.
Style: Warm tungsten, soft film grain, muted beige and gold palette.
Audio: Quiet room ambience, single soft water-drop sound at 2s, no dialogue.
Negative: No text, no logos, no harsh shadows, no reflections of crew.

Template 2 — Lipstick application beauty shot

Subject: A matte lipstick in a deep burgundy shade being applied to lips.
Action: The lipstick glides across the lower lip in one slow, deliberate stroke.
Scene: Close-up of a model's lips and chin, soft pink backdrop out of focus.
Camera: Static medium close-up, eye level, slight handheld breath.
Lens: 85mm, shallow depth of field.
Style: Soft beauty-dish key light, petal diffuser, pastel color grade.
Audio: Silence except for a faint breath inhale at 1s and 4s.
Negative: No teeth visible, no text overlays, no skin blemishes.

Template 3 — Skincare routine flat lay

Subject: Five skincare bottles and jars arranged in a row on a stone tray.
Action: Morning light sweeps across the products from left to right as the camera orbits.
Scene: Bathroom counter, eucalyptus sprig to the right, white linen towel beneath.
Camera: Slow 180-degree orbit, medium shot.
Lens: 50mm, moderate depth of field.
Style: Spa-grade, clean whites, sage green accents, soft natural light.
Audio: Distant birdsong through an open window, soft ceramic clink at 3s.
Negative: No visible brand logos, no text, no people.

Tech gadgets

Template 4 — Wireless earbuds unboxing feel

Subject: A pair of matte-silver wireless earbuds in an open charging case.
Action: The case lid opens slowly under its own weight, revealing the earbuds inside.
Scene: Dark walnut desk, single desk lamp casting warm pool of light.
Camera: Slow dolly-in, high angle 30 degrees.
Lens: 50mm, shallow depth of field, f/2.0.
Style: Premium product commercial, desaturated teal and amber grade.
Audio: Soft mechanical hinge sound at 0.5s, quiet room tone, no music.
Negative: No charging cables, no people, no packaging, no logos.

Template 5 — Smart watch on wrist

Subject: A titanium-cased smartwatch on a male wrist, screen glowing softly.
Action: The wrist rotates 30 degrees, catching the light, screen displays a heart-rate pulse.
Scene: Mountain trail at sunrise, blurred pines in background.
Camera: Tracking alongside, medium close-up.
Lens: 85mm, shallow depth of field, f/2.8.
Style: Adventure-grade, golden hour, cinematic film emulation.
Audio: Soft wind, distant raven call at 3s, no dialogue.
Negative: No visible screen text, no other people, no harsh shadows.

Template 6 — Mechanical keyboard typing

Subject: A low-profile mechanical keyboard with RGB backlighting under each key.
Action: Hands type a short phrase in real time, keys depress with tactile precision.
Scene: Dark desk setup, single monitor glow in background, blue ambient.
Camera: Overhead 45-degree angle, static.
Lens: 35mm, deep depth of field.
Style: Cyberpunk-influenced, neon blues and purples, high contrast.
Audio: Crisp mechanical key clicks, soft hum of computer fans, no music.
Negative: No faces visible, no screen text legible, no logos.

Food and beverage

Template 7 — Coffee pour hero

Subject: A cup of black coffee in a ceramic mug, viewed from a 30-degree angle.
Action: Steaming milk is poured in a thin stream, creating latte art.
Scene: Rustic wooden table, morning window light from camera left.
Camera: Static medium close-up.
Lens: 50mm, shallow depth of field, f/1.8.
Style: Warm, inviting, muted earth tones, slight vignette.
Audio: Soft pour of liquid, gentle ceramic scrape at 5s, distant kettle whistle.
Negative: No visible brand names, no people, no harsh reflections.

Template 8 — Craft beer bottle condensation

Subject: A chilled craft beer bottle with visible condensation droplets.
Action: A hand enters frame and rotates the bottle 45 degrees to reveal the label.
Scene: Dark wood bar, amber backlight, glassware in soft focus behind.
Camera: Medium close-up, slow push-in.
Lens: 85mm, very shallow depth of field.
Style: Moody pub atmosphere, warm tungsten, cinematic grain.
Audio: Ice clinking in a distant glass, faint jazz piano two rooms over.
Negative: No legible label text, no visible faces, no neon signage.

Template 9 — Chocolate truffle reveal

Subject: A single dark chocolate truffle dusted with cocoa powder.
Action: A hand places the truffle on a slate plate next to two more truffles.
Scene: Velvet tablecloth, candlelight from a single taper to the right.
Camera: Slow orbit, macro to medium transition.
Lens: Macro 100mm equivalent, extreme shallow depth of field.
Style: Rich, low-key lighting, deep browns and golds, soft glow.
Audio: Soft cloth rustle, single breath at 2s, no music.
Negative: No packaging, no people visible past the hand, no text.

Fashion and luxury

Template 10 — Silk scarf motion

Subject: A patterned silk scarf in jewel tones, 90cm square.
Action: The scarf drifts through still air, twisting into an S-curve.
Scene: Pure black void with a single soft spotlight from above.
Camera: Static medium shot, centered.
Lens: 85mm, moderate depth of field.
Style: High-fashion editorial, saturated color, slow-motion feel.
Audio: Silence, single breath of wind, no music, no dialogue.
Negative: No visible hands holding the scarf, no wrinkles, no text.

Template 11 — Leather handbag detail

Subject: A caramel-brown leather handbag with brass hardware.
Action: The bag is rotated slowly to show stitching and clasp detail.
Scene: Marble pedestal, museum-style spotlight, white seamless background.
Camera: Slow orbit, medium close-up.
Lens: 85mm, shallow depth of field.
Style: Luxury catalog, neutral palette, soft contrast, fine grain.
Audio: Single soft leather creak at 3s, otherwise silent.
Negative: No price tags, no people, no visible brand marks, no harsh shadows.

Template 12 — Perfume bottle hero

Subject: A faceted crystal perfume bottle with a gold atomizer.
Action: A fine mist sprays once to the right, catching the backlight.
Scene: Velvet dressing table, single candle to the left, mirror reflection behind.
Camera: Static, eye-level, medium close-up.
Lens: 100mm macro equivalent, very shallow depth of field.
Style: Glamour, deep contrast, ruby and gold tones, soft bloom.
Audio: Soft hiss of the atomizer at 1s, otherwise silent.
Negative: No visible text or labels, no reflections of crew, no harsh flash.

Lifestyle context

Template 13 — Yoga mat morning routine

Subject: A rolled natural-rubber yoga mat on a wooden floor at sunrise.
Action: A pair of hands unrolls the mat smoothly across the frame.
Scene: Sunlit apartment, large window with sheer curtains, plants in background.
Camera: Low angle, tracking the unroll motion.
Lens: 35mm, moderate depth of field.
Style: Wellness editorial, soft golden tones, film emulation.
Audio: Quiet morning ambience, distant traffic murmur, no music, no dialogue.
Negative: No visible people past hands, no logos, no clutter.

Template 14 — Running shoe on trail

Subject: A trail running shoe in burnt orange, planted on a rocky path.
Action: The shoe pushes off, kicking up a small spray of dust and gravel.
Scene: Mountain ridge at dawn, fog in the valley below, pines on the horizon.
Camera: Low angle, tracking alongside, 60fps slow-motion feel.
Lens: 35mm, deep depth of field.
Style: Adventure documentary, cool teals and warm orange contrast.
Audio: Crisp footfall on gravel, soft wind, single raven call at 4s.
Negative: No people visible past ankle, no logos readable, no motion blur.

Template 15 — Pour-over coffee ritual

Subject: A glass pour-over coffee carafe on a copper stand.
Action: Hot water pours in a slow spiral over a bed of ground coffee, blooming.
Scene: Morning kitchen, wood counters, plants on the windowsill.
Camera: Overhead 90-degree angle, static.
Lens: 50mm, moderate depth of field.
Style: Slow living, soft natural light, pastel beige palette.
Audio: Soft pour of water, faint percolation, distant birdsong, no dialogue.
Negative: No hands visible, no logos, no harsh reflections on glass.

4K output settings and aspect ratios

Resolution

Veo 3.1 supports multiple output resolutions. The 4K target is 3840×2160, the broadcast UHD standard. Lower resolutions (1080p, 720p) are available but defeat the purpose of using Veo 3.1 for premium product work.

In Vertex AI, select 4K explicitly in the generation parameters; the Gemini API default may be lower. The Google AI Studio Veo prompt cookbook lists the supported resolutions per region and account type.

Aspect ratios

Aspect Resolution (4K) Use case
16:9 3840×2160 YouTube, broadcast, web hero banners
9:16 2160×3840 TikTok, Reels, Shorts, Stories
1:1 2160×2160 Instagram feed, PDP galleries
4:5 2160×2700 Instagram feed portrait, Pinterest
21:9 3840×1620 Cinematic letterbox, hero website

Specify the aspect in your prompt or in the API call. Veo 3.1 does not always infer aspect from the prompt alone.

Generation time and cost

Veo 3.1 generations at 4K take longer and cost more than 1080p. Plan prompt iterations at 1080p first, then re-generate at 4K only on the final approved prompt. This is the standard workflow in the Google AI Studio Veo prompt cookbook.


Avoiding Veo 3.1’s weak spots

Duration control

Veo 3.1’s default clip length is 8 seconds. Longer clips require either stitching generations or using the model’s extended-clip mode where available. If you need a 15-second hero clip, plan two 8-second generations with a visual continuity cue in the second prompt.

Tip: end prompt 1 with a clear “ending pose” (a hand resting, a bottle stationary) so prompt 2 can resume naturally.

Multi-character complexity

Two on-screen speakers with synchronized dialogue is the single most failure-prone scenario in Veo 3.1. The model can produce two characters visually, but lip-sync to specific lines often mismatches. Workarounds:

  • Use one speaker at a time, cutting between them.
  • Use off-screen voiceover instead of on-screen dialogue.
  • Use the model for visual only, then add the dialogue in post.

Text and logos

Veo 3.1 still struggles with legible on-screen text. Brand logos frequently appear as blurry approximations. Do not ask Veo to render your wordmark. Add it in post.

Hand and finger artifacts

As with most current video models, hands can show extra or merged fingers. The negative prompt “no extra fingers, no deformed hands” helps but is not always honored. Plan a manual review pass for any clip where hands are foregrounded.

Consistency across multiple generations

For a campaign requiring the same product in five scenes, run one “hero” generation first, then describe the resulting visual in subsequent prompts (“the same matte-black bottle from the hero shot, now on a wooden shelf”). This reference-back technique is documented in the Veo prompt cookbook.


Frequently asked questions

What is Veo 3.1’s maximum clip length? Veo 3.1 generates 8-second clips by default. Extended modes may produce longer clips depending on your Vertex AI or Gemini API tier.

Does Veo 3.1 really generate audio natively? Yes. Dialogue, sound effects, and ambient sound are produced during the same inference pass as the video. No post-dubbing is required.

Can Veo 3.1 output 4K? Yes, up to 3840×2160 at supported aspect ratios. You must select 4K explicitly in your generation parameters.

What aspect ratios does Veo 3.1 support? 16:9, 9:16, 1:1, 4:5, and 21:9 are commonly supported. Verify in your specific Vertex AI region and account configuration.

How does Veo 3.1 compare to Sora 2 for product video? A November 2025 model comparison: Seedance vs Kling vs Veo notes Veo 3.1’s stronger native audio fidelity. Sora 2 tends to produce longer default clips but with weaker synchronized dialogue.

Can I use Veo 3.1 output commercially? Yes, under the standard Google generative AI terms. Always verify the latest commercial-use policy in your Vertex AI agreement.

Where can I find more product video prompt examples? Google AI Studio Veo prompt cookbook maintains a 25-prompt product library, and Google AI Studio Veo prompt cookbook curates commercial-focused prompts.

How do I keep a product consistent across multiple clips? Use the reference-back technique: generate one hero shot, then describe the resulting visual in subsequent prompts. The Veo prompt cookbook covers this in detail.


Conclusion

Veo 3.1’s combination of native audio and 4K output makes it the strongest single-prompt product video model available in late 2025. The model rewards cinematographer-grade vocabulary, structured prompts, and explicit audio layering. Use the templates above as starting points — substitute your product, your color palette, and your SFX, and iterate at 1080p before committing to a 4K generation.

For related reading on the broader AI video landscape, see our guides on Seedance 2.5 character motion prompts, Kling 3.0 character consistency workflows, and Wan 3.0 commercial ad pipelines. For a broader prompt-engineering foundation, the best AI video prompts 2026 masterclass and the 2K video model template library are worth bookmarking.

The official Google DeepMind Veo model page and the Google AI Studio Veo prompt cookbook remain the canonical references — bookmark them, and revisit as the model iterates.


Reviewed by the videosprompt.org editorial team · October 2026

Share Article

Related Articles

Recommended Reading

Ready to Get Started?

Experience our product immediately and explore more possibilities.