VideosPrompt VideosPrompt

Veo 3.1 原生音频与 4K 产品视频提示词:完整指南

作者: VideosPrompt 日期: 2026-10-07 14:11:55
Veo 3.1 原生音频与 4K 产品视频提示词:完整指南

摘要

  • Veo 3.1 是 Google DeepMind 的旗舰视频模型,可在单次生成中同时输出原生音频(对话、音效、环境音)和 4K 分辨率视频——无需后期配音(deepmind.google/models/veo)。
  • 最可靠的 Veo 3.1 提示词公式是 主体 → 动作 → 场景 → 镜头 → 镜头语言 → 风格 → 音频 → 负面提示,需严格按照此顺序排列。
  • Veo 3.1 对摄影师级专业词汇响应强烈——”低角度”、”变焦对焦”、”向左跟踪”、”35mm 镜头”和”伦勃朗光”都会显著改变输出效果(aistudio.google.com/docs/veo/cookbook)。
  • 原生音频包含三个独立层(对话、音效、环境音)。每层都必须单独指定,否则 Veo 的生成结果会不一致。
  • 4K(3840×2160)输出需要在 Vertex AI 或 Gemini API 中明确选择;并非所有界面都默认显示该选项(Google AI Studio Veo prompt cookbook)。
  • Veo 3.1 已知的薄弱环节是时长控制(默认 8 秒片段)和多角色同步对话——下文故障排查部分均有应对方案。

Veo 3.1 是什么,以及为什么原生音频 + 4K 至关重要

Veo 3.1 是 Google DeepMind 发布的第三代视频基础模型,于 2025 年 10 月至 11 月期间陆续上线 Gemini API、Vertex AI 和 Flow。与前代不同,Veo 3.1 能从单条文本提示词中一次性生成同步音频和 4K 分辨率视频。相关文档可参阅官方 Veo 模型页面 和实用的 Veo 提示词手册。

对产品营销人员而言,三项能力改变了工作流程:

  1. 原生音频生成。 对话、音效和环境声在视频推理的同一过程中生成。你不再需要在 Runway、ElevenLabs 或 DAW 中进行唇形同步。
  2. 4K 输出。 Veo 3.1 支持在多种宽高比下进行 4K(3840×2160)生成,适用于网页主横幅、广播剪辑和高端电商详情页。
  3. 摄影师感知型提示词。 该模型基于影视术语训练,能够理解镜头、灯光和运镜词汇并做出有意义的响应。

2025 年 11 月的 model comparison: Seedance vs Kling vs Veo 将 Veo 3.1 与 OpenAI 的 Sora 2 进行了对比,指出 Veo 的音频保真度更强,而 Sora 2 的默认片段长度更长。对于产品制作而言,Veo 的音频层是决定性因素。

本指南是一份实战参考文档,涵盖了提示词公式、摄影师专业词汇、三层音频模型、跨五个产品类别的 15 个生产就绪模板,以及那些反复浪费额度的常见失败模式。


Veo 3.1 提示词公式

产品视频提示词最可靠的结构是一条八段式链。顺序很重要,因为 Veo 按位置对词元加权——靠前的词元主导构图,靠后的词元起修饰作用。

Subject → Action → Scene → Camera → Lens → Style → Audio → Negative
插槽 用途 示例
主体 主角产品或模特 “一只哑光黑色陶瓷护肤瓶”
动作 片段中发生的事 “缓慢旋转 90 度”
场景 所处环境 “在森林中一块湿润的石板上”
镜头 运动和构图 “缓慢推入,平视”
镜头语言 焦距和光圈表现 “85mm,浅景深”
风格 调色与视觉风格 “暖色钨丝灯,35mm 胶片颗粒”
音频 对话、音效、环境音 “轻柔的雨声环境音,无对话,第 2 秒有一滴水滴音效”
负面提示 需要避免的内容 “不要文字、不要标志、不要多余手指、不要硬阴影”

Vertex AI 电商提示词库(Google AI Studio Veo prompt cookbook)也采用相同的结构,只是将镜头语言和风格合并为一个”外观”字段。将其拆分可获得更精细的控制。


使用摄影师的语言

Veo 3.1 的训练数据中包含了大量摄影参考资料,因此专业术语可以产生效果。模糊的表述(如”一个不错的镜头角度”)会降低输出质量,而具体的表述则会显著提升效果。

可靠有效的镜头运动

  • 推近 / 拉远 — 整机前后移动。适合产品揭幕。
  • 向左跟踪 / 向右跟踪 — 与主体平行的横向运动。适合行走镜头和传送带展示。
  • 升臂 / 降臂 — 沿垂直轴的升降运动。适合规模感镜头。
  • 低角度 / 高角度 — 视角指定。低角度让产品显得宏伟;高角度则更具编辑感。
  • 变焦对焦(Rack focus) — 在前后景之间切换焦平面。Veo 3.1 对静态构图处理良好,但在快速运动中表现欠佳。
  • 手持 — 增加可控的抖动。在生活方式内容中谨慎使用;如果提示词同时要求”电影感”,可能显得业余。
  • 环绕 / 弧线运动 — 围绕主体的圆周运动。产品主镜头最可靠的运动方式。

构图术语

  • 特写(CU)、中特写(MCU)、中景(MS)、全景(WS) — 标准景别。
  • 过肩镜头(OTS) — 用于带主持人的对话或产品演示。
  • 插入镜头 / 微距 — 极近距离的特写,用于呈现质感、成分或细节叙事。
  • 双人镜头 — 用于产品与人物或成对物品同时出现的场景。

Veo 提示词手册 证实,将运动方式与构图术语结合使用(例如”缓慢环绕,中特写”)比单独使用任何一项都能获得更稳定的结果。


原生音频:对话、音效、环境音

Veo 3.1 的音频模型包含三个独立层。在提示词中分别指定每一层——而不是笼统地写”音效”——是产出可用渲染和音频错位/缺失片段之间的关键差别。

第一层:对话

将台词用引号标注并明确放置:

Audio: A female voiceover says, "Meet the new Hydra-Serum. Three drops, every morning."

对于同步的画外音台词,请注明说话者并给出唇形提示:

Audio: A male presenter, off-camera, says, "Notice the brushed-aluminum finish."

Veo 3.1 能够可靠地处理单说话者对话。双说话者同步对话是其已知薄弱环节之一(见下文)。

第二层:音效(SFX)

要具体。”飞溅”只产生一个通用飞溅声。”第 1 秒一滴水击中大理石”则能产生一个精确的电影级音效。

常用有效的音效表述方式:

  • “磁吸闭合的柔和机械咔嗒声”
  • “第 2 秒单次浓缩咖啡萃取,奶泡嘶嘶声”
  • “轻微的纸张撕裂声”
  • “玻璃杯轻轻放在杯垫上”

第三层:环境音

环境音是基础层——室内音、风声、远处的车流声。缺少它,片段会显得空洞。

  • “安静的森林环境音,远处鸟鸣,无风声”
  • “黄金时段的露天屋顶露台,隐约的城市嗡鸣”
  • “摄影棚内景,死寂声学环境,无混响”

Google AI Studio Veo prompt cookbook 指出,明确指定环境音的产出音频明显比依赖 Veo 默认行为更加”制作精良”。


有效的镜头与灯光词汇

灯光方案

术语 效果 适用场景
伦勃朗光 眼下形成三角形光斑,经典人像布光 护肤、香水、主形象人像
分割光 半张脸打亮,半张脸处于阴影中 神秘、高端、男性护理
蝴蝶光 / 派拉蒙光 来自上方的光线,鼻下形成小阴影 美妆、华丽、珠宝
轮廓光 / 逆光 主体边缘的高光 将产品与背景分离
黄金时段 温暖的低位阳光,长长的阴影 户外生活方式、时尚
阴天柔光 漫射光,无阴影 护肤微距、食物
霓虹光 / 实用光 来自画面内光源的色彩 科技、游戏、夜生活

镜头与光圈

  • 35mm — 标准广角,带轻度环境感。多用途。
  • 50mm — “自然”焦距,几乎无畸变。编辑类产品的默认选择。
  • 85mm — 讨喜的压缩感,经典人像焦距。美妆主镜头的默认选择。
  • 微距 — 极近距离特写,浅景深,展现质感。
  • 变形宽银幕 — 椭圆形散景,水平光斑。电影感强但可能扭曲产品几何形状。

请始终将镜头与光圈表现配对使用:

Lens: 85mm, shallow depth of field, f/1.8 equivalent

不指定光圈时,Veo 的默认值不稳定。


15 个生产就绪的产品视频模板

以下每个模板都格式化为一条完整提示词。可原样复制,然后替换你的产品、品牌颜色或特定音效。

美妆与护肤

模板 1 — 精华液滴落主镜头

Subject: A 30ml matte-black glass serum bottle with a gold dropper cap.
Action: A single amber drop falls from the dropper onto the back of a hand.
Scene: On a white marble slab, soft window light from camera left, dust motes visible.
Camera: Slow dolly-in, macro framing, eye level.
Lens: 85mm, shallow depth of field, f/2.0.
Style: Warm tungsten, soft film grain, muted beige and gold palette.
Audio: Quiet room ambience, single soft water-drop sound at 2s, no dialogue.
Negative: No text, no logos, no harsh shadows, no reflections of crew.

模板 2 — 口红涂抹美妆镜头

Subject: A matte lipstick in a deep burgundy shade being applied to lips.
Action: The lipstick glides across the lower lip in one slow, deliberate stroke.
Scene: Close-up of a model's lips and chin, soft pink backdrop out of focus.
Camera: Static medium close-up, eye level, slight handheld breath.
Lens: 85mm, shallow depth of field.
Style: Soft beauty-dish key light, petal diffuser, pastel color grade.
Audio: Silence except for a faint breath inhale at 1s and 4s.
Negative: No teeth visible, no text overlays, no skin blemishes.

模板 3 — 护肤流程平铺俯拍

Subject: Five skincare bottles and jars arranged in a row on a stone tray.
Action: Morning light sweeps across the products from left to right as the camera orbits.
Scene: Bathroom counter, eucalyptus sprig to the right, white linen towel beneath.
Camera: Slow 180-degree orbit, medium shot.
Lens: 50mm, moderate depth of field.
Style: Spa-grade, clean whites, sage green accents, soft natural light.
Audio: Distant birdsong through an open window, soft ceramic clink at 3s.
Negative: No visible brand logos, no text, no people.

科技数码

模板 4 — 无线耳机开箱质感

Subject: A pair of matte-silver wireless earbuds in an open charging case.
Action: The case lid opens slowly under its own weight, revealing the earbuds inside.
Scene: Dark walnut desk, single desk lamp casting warm pool of light.
Camera: Slow dolly-in, high angle 30 degrees.
Lens: 50mm, shallow depth of field, f/2.0.
Style: Premium product commercial, desaturated teal and amber grade.
Audio: Soft mechanical hinge sound at 0.5s, quiet room tone, no music.
Negative: No charging cables, no people, no packaging, no logos.

模板 5 — 腕上的智能手表

Subject: A titanium-cased smartwatch on a male wrist, screen glowing softly.
Action: The wrist rotates 30 degrees, catching the light, screen displays a heart-rate pulse.
Scene: Mountain trail at sunrise, blurred pines in background.
Camera: Tracking alongside, medium close-up.
Lens: 85mm, shallow depth of field, f/2.8.
Style: Adventure-grade, golden hour, cinematic film emulation.
Audio: Soft wind, distant raven call at 3s, no dialogue.
Negative: No visible screen text, no other people, no harsh shadows.

模板 6 — 机械键盘敲击

Subject: A low-profile mechanical keyboard with RGB backlighting under each key.
Action: Hands type a short phrase in real time, keys depress with tactile precision.
Scene: Dark desk setup, single monitor glow in background, blue ambient.
Camera: Overhead 45-degree angle, static.
Lens: 35mm, deep depth of field.
Style: Cyberpunk-influenced, neon blues and purples, high contrast.
Audio: Crisp mechanical key clicks, soft hum of computer fans, no music.
Negative: No faces visible, no screen text legible, no logos.

食品与饮料

模板 7 — 咖啡倾倒主镜头

Subject: A cup of black coffee in a ceramic mug, viewed from a 30-degree angle.
Action: Steaming milk is poured in a thin stream, creating latte art.
Scene: Rustic wooden table, morning window light from camera left.
Camera: Static medium close-up.
Lens: 50mm, shallow depth of field, f/1.8.
Style: Warm, inviting, muted earth tones, slight vignette.
Audio: Soft pour of liquid, gentle ceramic scrape at 5s, distant kettle whistle.
Negative: No visible brand names, no people, no harsh reflections.

模板 8 — 精酿啤酒瓶凝露

Subject: A chilled craft beer bottle with visible condensation droplets.
Action: A hand enters frame and rotates the bottle 45 degrees to reveal the label.
Scene: Dark wood bar, amber backlight, glassware in soft focus behind.
Camera: Medium close-up, slow push-in.
Lens: 85mm, very shallow depth of field.
Style: Moody pub atmosphere, warm tungsten, cinematic grain.
Audio: Ice clinking in a distant glass, faint jazz piano two rooms over.
Negative: No legible label text, no visible faces, no neon signage.

模板 9 — 松露巧克力展示

Subject: A single dark chocolate truffle dusted with cocoa powder.
Action: A hand places the truffle on a slate plate next to two more truffles.
Scene: Velvet tablecloth, candlelight from a single taper to the right.
Camera: Slow orbit, macro to medium transition.
Lens: Macro 100mm equivalent, extreme shallow depth of field.
Style: Rich, low-key lighting, deep browns and golds, soft glow.
Audio: Soft cloth rustle, single breath at 2s, no music.
Negative: No packaging, no people visible past the hand, no text.

时尚与奢侈品

模板 10 — 真丝围巾飘动

Subject: A patterned silk scarf in jewel tones, 90cm square.
Action: The scarf drifts through still air, twisting into an S-curve.
Scene: Pure black void with a single soft spotlight from above.
Camera: Static medium shot, centered.
Lens: 85mm, moderate depth of field.
Style: High-fashion editorial, saturated color, slow-motion feel.
Audio: Silence, single breath of wind, no music, no dialogue.
Negative: No visible hands holding the scarf, no wrinkles, no text.

模板 11 — 皮质手袋细节

Subject: A caramel-brown leather handbag with brass hardware.
Action: The bag is rotated slowly to show stitching and clasp detail.
Scene: Marble pedestal, museum-style spotlight, white seamless background.
Camera: Slow orbit, medium close-up.
Lens: 85mm, shallow depth of field.
Style: Luxury catalog, neutral palette, soft contrast, fine grain.
Audio: Single soft leather creak at 3s, otherwise silent.
Negative: No price tags, no people, no visible brand marks, no harsh shadows.

模板 12 — 香水瓶主镜头

Subject: A faceted crystal perfume bottle with a gold atomizer.
Action: A fine mist sprays once to the right, catching the backlight.
Scene: Velvet dressing table, single candle to the left, mirror reflection behind.
Camera: Static, eye-level, medium close-up.
Lens: 100mm macro equivalent, very shallow depth of field.
Style: Glamour, deep contrast, ruby and gold tones, soft bloom.
Audio: Soft hiss of the atomizer at 1s, otherwise silent.
Negative: No visible text or labels, no reflections of crew, no harsh flash.

生活方式场景

模板 13 — 清晨瑜伽垫

Subject: A rolled natural-rubber yoga mat on a wooden floor at sunrise.
Action: A pair of hands unrolls the mat smoothly across the frame.
Scene: Sunlit apartment, large window with sheer curtains, plants in background.
Camera: Low angle, tracking the unroll motion.
Lens: 35mm, moderate depth of field.
Style: Wellness editorial, soft golden tones, film emulation.
Audio: Quiet morning ambience, distant traffic murmur, no music, no dialogue.
Negative: No visible people past hands, no logos, no clutter.

模板 14 — 越野跑鞋登山

Subject: A trail running shoe in burnt orange, planted on a rocky path.
Action: The shoe pushes off, kicking up a small spray of dust and gravel.
Scene: Mountain ridge at dawn, fog in the valley below, pines on the horizon.
Camera: Low angle, tracking alongside, 60fps slow-motion feel.
Lens: 35mm, deep depth of field.
Style: Adventure documentary, cool teals and warm orange contrast.
Audio: Crisp footfall on gravel, soft wind, single raven call at 4s.
Negative: No people visible past ankle, no logos readable, no motion blur.

模板 15 — 手冲咖啡仪式

Subject: A glass pour-over coffee carafe on a copper stand.
Action: Hot water pours in a slow spiral over a bed of ground coffee, blooming.
Scene: Morning kitchen, wood counters, plants on the windowsill.
Camera: Overhead 90-degree angle, static.
Lens: 50mm, moderate depth of field.
Style: Slow living, soft natural light, pastel beige palette.
Audio: Soft pour of water, faint percolation, distant birdsong, no dialogue.
Negative: No hands visible, no logos, no harsh reflections on glass.

4K 输出设置与宽高比

分辨率

Veo 3.1 支持多种输出分辨率。4K 目标为 3840×2160,即广播级 UHD 标准。更低分辨率(1080p、720p)也可使用,但无法发挥 Veo 3.1 在高端产品制作中的优势。

在 Vertex AI 中,需在生成参数中明确选择 4K;Gemini API 的默认分辨率可能较低。Google AI Studio Veo prompt cookbook 列出了各地区和账户类型支持的分辨率。

宽高比

宽高比 分辨率(4K) 适用场景
16:9 3840×2160 YouTube、广播、网页主横幅
9:16 2160×3840 TikTok、Reels、Shorts、Stories
1:1 2160×2160 Instagram 信息流、PDP 画廊
4:5 2160×2700 Instagram 信息流竖图、Pinterest
21:9 3840×1620 电影信箱模式、高端网站主视觉

在提示词或 API 调用中指定宽高比。Veo 3.1 并不总是能从提示词中自动推断宽高比。

生成时间与成本

Veo 3.1 的 4K 生成比 1080p 更耗时且成本更高。建议先用 1080p 进行提示词迭代,最后在确认通过的提示词上以 4K 重新生成。这正是 Google AI Studio Veo prompt cookbook 中推荐的标准工作流。


规避 Veo 3.1 的薄弱环节

时长控制

Veo 3.1 的默认片段长度为 8 秒。更长的片段需要拼接生成结果,或在可用时使用模型的扩展片段模式。如果需要 15 秒的主镜头,建议生成两段 8 秒片段,并在第二段提示词中加入视觉连续性提示。

技巧:在第一段提示词结尾设置一个清晰的”结束姿态”(如手部停放、产品静止),以便第二段自然衔接。

多角色复杂度

两位画内说话者的同步对话是 Veo 3.1 中最容易失败的场景。模型可以视觉上生成两个角色,但唇形与特定台词的同步常常失配。变通方案:

  • 一次只使用一个说话者,通过剪辑切换。
  • 使用画外旁白代替画内对话。
  • 仅使用模型生成画面,对话在后期添加。

文字与标志

Veo 3.1 仍然难以生成清晰的画面内文字。品牌标志经常呈现为模糊的近似形态。不要让 Veo 渲染你的文字商标,请在后期添加。

手部与手指瑕疵

与大多数当前视频模型一样,手部可能出现多余或合并的手指。负面提示”不要多余手指,不要变形的手”有助于改善,但并非总能生效。如果片段中手部是前景元素,建议安排人工审核。

多片段之间的一致性

对于需要在五个场景中展示同一产品的推广活动,建议先生成一段”主形象”片段,然后在后续提示词中描述该视觉结果(”与主形象镜头相同的哑光黑色瓶子,现在置于木质架子上”)。这种回溯参考技术在 Veo 提示词手册 中有详细记录。


常见问题

Veo 3.1 的最大片段长度是多少? Veo 3.1 默认生成 8 秒片段。扩展模式可根据你的 Vertex AI 或 Gemini API 等级生成更长片段。

Veo 3.1 真的能原生生成音频吗? 是的。对话、音效和环境声在视频推理的同一过程中生成,无需后期配音。

Veo 3.1 能输出 4K 吗? 可以,在支持的宽高比下最高可达 3840×2160。你必须在生成参数中明确选择 4K。

Veo 3.1 支持哪些宽高比? 16:9、9:16、1:1、4:5 和 21:9 为常见支持的比例。具体请根据你的 Vertex AI 地区和账户配置核实。

在产品视频方面,Veo 3.1 与 Sora 2 相比如何? 2025 年 11 月的 model comparison: Seedance vs Kling vs Veo 指出 Veo 3.1 的原生音频保真度更强。Sora 2 的默认片段往往更长,但同步对话效果较弱。

我可以将 Veo 3.1 的输出用于商业用途吗? 可以,依据 Google 生成式 AI 的标准条款。请务必在你签署的 Vertex AI 协议中核实最新的商业使用政策。

在哪里可以找到更多产品视频提示词示例? Google AI Studio Veo prompt cookbook 维护着一个包含 25 条提示词的产品库,Google AI Studio Veo prompt cookbook 则整理了面向商业用途的提示词。

如何让一个产品在不同片段之间保持一致? 使用回溯参考技术:先生成一段主形象镜头,然后在后续提示词中描述该视觉结果。Veo 提示词手册 对此有详细阐述。


结语

Veo 3.1 将原生音频和 4K 输出相结合,使其成为 2025 年末最强大的单提示词产品视频模型。该模型对摄影师级词汇、结构化提示词和明确音频分层均有良好响应。以上模板可作为起点——替换你的产品、调色板和音效,先在 1080p 下迭代,最后再提交 4K 生成。

如需了解更广泛的 AI 视频领域相关内容,请参阅我们关于 Seedance 2.5 角色动作提示词、Kling 3.0 角色一致性工作流 和 Wan 3.0 商业广告流水线 的指南。如需打下更扎实的提示词工程基础,2026 年最佳 AI 视频提示词大师课 和 2K 视频模型模板库 值得收藏。

官方 Google DeepMind Veo 模型页面 和 Google AI Studio Veo prompt cookbook 仍然是权威参考——请将它们加入书签,并在模型迭代时回顾查看。


由 videosprompt.org 编辑团队审校 · 2026 年 10 月

分享文章

相关文章

推荐阅读

开始你的下一步

探索更多可能,发现适合你的解决方案。