Veo 3.1 原生音频与 4K 产品视频提示词:完整指南
摘要
- Veo 3.1 是 Google DeepMind 的旗舰视频模型,可在单次生成中同时输出原生音频(对话、音效、环境音)和 4K 分辨率视频——无需后期配音(deepmind.google/models/veo)。
- 最可靠的 Veo 3.1 提示词公式是
主体 → 动作 → 场景 → 镜头 → 镜头语言 → 风格 → 音频 → 负面提示,需严格按照此顺序排列。 - Veo 3.1 对摄影师级专业词汇响应强烈——”低角度”、”变焦对焦”、”向左跟踪”、”35mm 镜头”和”伦勃朗光”都会显著改变输出效果(aistudio.google.com/docs/veo/cookbook)。
- 原生音频包含三个独立层(对话、音效、环境音)。每层都必须单独指定,否则 Veo 的生成结果会不一致。
- 4K(3840×2160)输出需要在 Vertex AI 或 Gemini API 中明确选择;并非所有界面都默认显示该选项(Google AI Studio Veo prompt cookbook)。
- Veo 3.1 已知的薄弱环节是时长控制(默认 8 秒片段)和多角色同步对话——下文故障排查部分均有应对方案。
Veo 3.1 是什么,以及为什么原生音频 + 4K 至关重要
Veo 3.1 是 Google DeepMind 发布的第三代视频基础模型,于 2025 年 10 月至 11 月期间陆续上线 Gemini API、Vertex AI 和 Flow。与前代不同,Veo 3.1 能从单条文本提示词中一次性生成同步音频和 4K 分辨率视频。相关文档可参阅官方 Veo 模型页面 和实用的 Veo 提示词手册。
对产品营销人员而言,三项能力改变了工作流程:
- 原生音频生成。 对话、音效和环境声在视频推理的同一过程中生成。你不再需要在 Runway、ElevenLabs 或 DAW 中进行唇形同步。
- 4K 输出。 Veo 3.1 支持在多种宽高比下进行 4K(3840×2160)生成,适用于网页主横幅、广播剪辑和高端电商详情页。
- 摄影师感知型提示词。 该模型基于影视术语训练,能够理解镜头、灯光和运镜词汇并做出有意义的响应。
2025 年 11 月的 model comparison: Seedance vs Kling vs Veo 将 Veo 3.1 与 OpenAI 的 Sora 2 进行了对比,指出 Veo 的音频保真度更强,而 Sora 2 的默认片段长度更长。对于产品制作而言,Veo 的音频层是决定性因素。
本指南是一份实战参考文档,涵盖了提示词公式、摄影师专业词汇、三层音频模型、跨五个产品类别的 15 个生产就绪模板,以及那些反复浪费额度的常见失败模式。
Veo 3.1 提示词公式
产品视频提示词最可靠的结构是一条八段式链。顺序很重要,因为 Veo 按位置对词元加权——靠前的词元主导构图,靠后的词元起修饰作用。
Subject → Action → Scene → Camera → Lens → Style → Audio → Negative
| 插槽 | 用途 | 示例 |
|---|---|---|
| 主体 | 主角产品或模特 | “一只哑光黑色陶瓷护肤瓶” |
| 动作 | 片段中发生的事 | “缓慢旋转 90 度” |
| 场景 | 所处环境 | “在森林中一块湿润的石板上” |
| 镜头 | 运动和构图 | “缓慢推入,平视” |
| 镜头语言 | 焦距和光圈表现 | “85mm,浅景深” |
| 风格 | 调色与视觉风格 | “暖色钨丝灯,35mm 胶片颗粒” |
| 音频 | 对话、音效、环境音 | “轻柔的雨声环境音,无对话,第 2 秒有一滴水滴音效” |
| 负面提示 | 需要避免的内容 | “不要文字、不要标志、不要多余手指、不要硬阴影” |
Vertex AI 电商提示词库(Google AI Studio Veo prompt cookbook)也采用相同的结构,只是将镜头语言和风格合并为一个”外观”字段。将其拆分可获得更精细的控制。
使用摄影师的语言
Veo 3.1 的训练数据中包含了大量摄影参考资料,因此专业术语可以产生效果。模糊的表述(如”一个不错的镜头角度”)会降低输出质量,而具体的表述则会显著提升效果。
可靠有效的镜头运动
- 推近 / 拉远 — 整机前后移动。适合产品揭幕。
- 向左跟踪 / 向右跟踪 — 与主体平行的横向运动。适合行走镜头和传送带展示。
- 升臂 / 降臂 — 沿垂直轴的升降运动。适合规模感镜头。
- 低角度 / 高角度 — 视角指定。低角度让产品显得宏伟;高角度则更具编辑感。
- 变焦对焦(Rack focus) — 在前后景之间切换焦平面。Veo 3.1 对静态构图处理良好,但在快速运动中表现欠佳。
- 手持 — 增加可控的抖动。在生活方式内容中谨慎使用;如果提示词同时要求”电影感”,可能显得业余。
- 环绕 / 弧线运动 — 围绕主体的圆周运动。产品主镜头最可靠的运动方式。
构图术语
- 特写(CU)、中特写(MCU)、中景(MS)、全景(WS) — 标准景别。
- 过肩镜头(OTS) — 用于带主持人的对话或产品演示。
- 插入镜头 / 微距 — 极近距离的特写,用于呈现质感、成分或细节叙事。
- 双人镜头 — 用于产品与人物或成对物品同时出现的场景。
Veo 提示词手册 证实,将运动方式与构图术语结合使用(例如”缓慢环绕,中特写”)比单独使用任何一项都能获得更稳定的结果。
原生音频:对话、音效、环境音
Veo 3.1 的音频模型包含三个独立层。在提示词中分别指定每一层——而不是笼统地写”音效”——是产出可用渲染和音频错位/缺失片段之间的关键差别。
第一层:对话
将台词用引号标注并明确放置:
Audio: A female voiceover says, "Meet the new Hydra-Serum. Three drops, every morning."
对于同步的画外音台词,请注明说话者并给出唇形提示:
Audio: A male presenter, off-camera, says, "Notice the brushed-aluminum finish."
Veo 3.1 能够可靠地处理单说话者对话。双说话者同步对话是其已知薄弱环节之一(见下文)。
第二层:音效(SFX)
要具体。”飞溅”只产生一个通用飞溅声。”第 1 秒一滴水击中大理石”则能产生一个精确的电影级音效。
常用有效的音效表述方式:
- “磁吸闭合的柔和机械咔嗒声”
- “第 2 秒单次浓缩咖啡萃取,奶泡嘶嘶声”
- “轻微的纸张撕裂声”
- “玻璃杯轻轻放在杯垫上”
第三层:环境音
环境音是基础层——室内音、风声、远处的车流声。缺少它,片段会显得空洞。
- “安静的森林环境音,远处鸟鸣,无风声”
- “黄金时段的露天屋顶露台,隐约的城市嗡鸣”
- “摄影棚内景,死寂声学环境,无混响”
Google AI Studio Veo prompt cookbook 指出,明确指定环境音的产出音频明显比依赖 Veo 默认行为更加”制作精良”。
有效的镜头与灯光词汇
灯光方案
| 术语 | 效果 | 适用场景 |
|---|---|---|
| 伦勃朗光 | 眼下形成三角形光斑,经典人像布光 | 护肤、香水、主形象人像 |
| 分割光 | 半张脸打亮,半张脸处于阴影中 | 神秘、高端、男性护理 |
| 蝴蝶光 / 派拉蒙光 | 来自上方的光线,鼻下形成小阴影 | 美妆、华丽、珠宝 |
| 轮廓光 / 逆光 | 主体边缘的高光 | 将产品与背景分离 |
| 黄金时段 | 温暖的低位阳光,长长的阴影 | 户外生活方式、时尚 |
| 阴天柔光 | 漫射光,无阴影 | 护肤微距、食物 |
| 霓虹光 / 实用光 | 来自画面内光源的色彩 | 科技、游戏、夜生活 |
镜头与光圈
- 35mm — 标准广角,带轻度环境感。多用途。
- 50mm — “自然”焦距,几乎无畸变。编辑类产品的默认选择。
- 85mm — 讨喜的压缩感,经典人像焦距。美妆主镜头的默认选择。
- 微距 — 极近距离特写,浅景深,展现质感。
- 变形宽银幕 — 椭圆形散景,水平光斑。电影感强但可能扭曲产品几何形状。
请始终将镜头与光圈表现配对使用:
Lens: 85mm, shallow depth of field, f/1.8 equivalent
不指定光圈时,Veo 的默认值不稳定。
15 个生产就绪的产品视频模板
以下每个模板都格式化为一条完整提示词。可原样复制,然后替换你的产品、品牌颜色或特定音效。
美妆与护肤
模板 1 — 精华液滴落主镜头
Subject: A 30ml matte-black glass serum bottle with a gold dropper cap.
Action: A single amber drop falls from the dropper onto the back of a hand.
Scene: On a white marble slab, soft window light from camera left, dust motes visible.
Camera: Slow dolly-in, macro framing, eye level.
Lens: 85mm, shallow depth of field, f/2.0.
Style: Warm tungsten, soft film grain, muted beige and gold palette.
Audio: Quiet room ambience, single soft water-drop sound at 2s, no dialogue.
Negative: No text, no logos, no harsh shadows, no reflections of crew.
模板 2 — 口红涂抹美妆镜头
Subject: A matte lipstick in a deep burgundy shade being applied to lips.
Action: The lipstick glides across the lower lip in one slow, deliberate stroke.
Scene: Close-up of a model's lips and chin, soft pink backdrop out of focus.
Camera: Static medium close-up, eye level, slight handheld breath.
Lens: 85mm, shallow depth of field.
Style: Soft beauty-dish key light, petal diffuser, pastel color grade.
Audio: Silence except for a faint breath inhale at 1s and 4s.
Negative: No teeth visible, no text overlays, no skin blemishes.
模板 3 — 护肤流程平铺俯拍
Subject: Five skincare bottles and jars arranged in a row on a stone tray.
Action: Morning light sweeps across the products from left to right as the camera orbits.
Scene: Bathroom counter, eucalyptus sprig to the right, white linen towel beneath.
Camera: Slow 180-degree orbit, medium shot.
Lens: 50mm, moderate depth of field.
Style: Spa-grade, clean whites, sage green accents, soft natural light.
Audio: Distant birdsong through an open window, soft ceramic clink at 3s.
Negative: No visible brand logos, no text, no people.
科技数码
模板 4 — 无线耳机开箱质感
Subject: A pair of matte-silver wireless earbuds in an open charging case.
Action: The case lid opens slowly under its own weight, revealing the earbuds inside.
Scene: Dark walnut desk, single desk lamp casting warm pool of light.
Camera: Slow dolly-in, high angle 30 degrees.
Lens: 50mm, shallow depth of field, f/2.0.
Style: Premium product commercial, desaturated teal and amber grade.
Audio: Soft mechanical hinge sound at 0.5s, quiet room tone, no music.
Negative: No charging cables, no people, no packaging, no logos.
模板 5 — 腕上的智能手表
Subject: A titanium-cased smartwatch on a male wrist, screen glowing softly.
Action: The wrist rotates 30 degrees, catching the light, screen displays a heart-rate pulse.
Scene: Mountain trail at sunrise, blurred pines in background.
Camera: Tracking alongside, medium close-up.
Lens: 85mm, shallow depth of field, f/2.8.
Style: Adventure-grade, golden hour, cinematic film emulation.
Audio: Soft wind, distant raven call at 3s, no dialogue.
Negative: No visible screen text, no other people, no harsh shadows.
模板 6 — 机械键盘敲击
Subject: A low-profile mechanical keyboard with RGB backlighting under each key.
Action: Hands type a short phrase in real time, keys depress with tactile precision.
Scene: Dark desk setup, single monitor glow in background, blue ambient.
Camera: Overhead 45-degree angle, static.
Lens: 35mm, deep depth of field.
Style: Cyberpunk-influenced, neon blues and purples, high contrast.
Audio: Crisp mechanical key clicks, soft hum of computer fans, no music.
Negative: No faces visible, no screen text legible, no logos.
食品与饮料
模板 7 — 咖啡倾倒主镜头
Subject: A cup of black coffee in a ceramic mug, viewed from a 30-degree angle.
Action: Steaming milk is poured in a thin stream, creating latte art.
Scene: Rustic wooden table, morning window light from camera left.
Camera: Static medium close-up.
Lens: 50mm, shallow depth of field, f/1.8.
Style: Warm, inviting, muted earth tones, slight vignette.
Audio: Soft pour of liquid, gentle ceramic scrape at 5s, distant kettle whistle.
Negative: No visible brand names, no people, no harsh reflections.
模板 8 — 精酿啤酒瓶凝露
Subject: A chilled craft beer bottle with visible condensation droplets.
Action: A hand enters frame and rotates the bottle 45 degrees to reveal the label.
Scene: Dark wood bar, amber backlight, glassware in soft focus behind.
Camera: Medium close-up, slow push-in.
Lens: 85mm, very shallow depth of field.
Style: Moody pub atmosphere, warm tungsten, cinematic grain.
Audio: Ice clinking in a distant glass, faint jazz piano two rooms over.
Negative: No legible label text, no visible faces, no neon signage.
模板 9 — 松露巧克力展示
Subject: A single dark chocolate truffle dusted with cocoa powder.
Action: A hand places the truffle on a slate plate next to two more truffles.
Scene: Velvet tablecloth, candlelight from a single taper to the right.
Camera: Slow orbit, macro to medium transition.
Lens: Macro 100mm equivalent, extreme shallow depth of field.
Style: Rich, low-key lighting, deep browns and golds, soft glow.
Audio: Soft cloth rustle, single breath at 2s, no music.
Negative: No packaging, no people visible past the hand, no text.
时尚与奢侈品
模板 10 — 真丝围巾飘动
Subject: A patterned silk scarf in jewel tones, 90cm square.
Action: The scarf drifts through still air, twisting into an S-curve.
Scene: Pure black void with a single soft spotlight from above.
Camera: Static medium shot, centered.
Lens: 85mm, moderate depth of field.
Style: High-fashion editorial, saturated color, slow-motion feel.
Audio: Silence, single breath of wind, no music, no dialogue.
Negative: No visible hands holding the scarf, no wrinkles, no text.
模板 11 — 皮质手袋细节
Subject: A caramel-brown leather handbag with brass hardware.
Action: The bag is rotated slowly to show stitching and clasp detail.
Scene: Marble pedestal, museum-style spotlight, white seamless background.
Camera: Slow orbit, medium close-up.
Lens: 85mm, shallow depth of field.
Style: Luxury catalog, neutral palette, soft contrast, fine grain.
Audio: Single soft leather creak at 3s, otherwise silent.
Negative: No price tags, no people, no visible brand marks, no harsh shadows.
模板 12 — 香水瓶主镜头
Subject: A faceted crystal perfume bottle with a gold atomizer.
Action: A fine mist sprays once to the right, catching the backlight.
Scene: Velvet dressing table, single candle to the left, mirror reflection behind.
Camera: Static, eye-level, medium close-up.
Lens: 100mm macro equivalent, very shallow depth of field.
Style: Glamour, deep contrast, ruby and gold tones, soft bloom.
Audio: Soft hiss of the atomizer at 1s, otherwise silent.
Negative: No visible text or labels, no reflections of crew, no harsh flash.
生活方式场景
模板 13 — 清晨瑜伽垫
Subject: A rolled natural-rubber yoga mat on a wooden floor at sunrise.
Action: A pair of hands unrolls the mat smoothly across the frame.
Scene: Sunlit apartment, large window with sheer curtains, plants in background.
Camera: Low angle, tracking the unroll motion.
Lens: 35mm, moderate depth of field.
Style: Wellness editorial, soft golden tones, film emulation.
Audio: Quiet morning ambience, distant traffic murmur, no music, no dialogue.
Negative: No visible people past hands, no logos, no clutter.
模板 14 — 越野跑鞋登山
Subject: A trail running shoe in burnt orange, planted on a rocky path.
Action: The shoe pushes off, kicking up a small spray of dust and gravel.
Scene: Mountain ridge at dawn, fog in the valley below, pines on the horizon.
Camera: Low angle, tracking alongside, 60fps slow-motion feel.
Lens: 35mm, deep depth of field.
Style: Adventure documentary, cool teals and warm orange contrast.
Audio: Crisp footfall on gravel, soft wind, single raven call at 4s.
Negative: No people visible past ankle, no logos readable, no motion blur.
模板 15 — 手冲咖啡仪式
Subject: A glass pour-over coffee carafe on a copper stand.
Action: Hot water pours in a slow spiral over a bed of ground coffee, blooming.
Scene: Morning kitchen, wood counters, plants on the windowsill.
Camera: Overhead 90-degree angle, static.
Lens: 50mm, moderate depth of field.
Style: Slow living, soft natural light, pastel beige palette.
Audio: Soft pour of water, faint percolation, distant birdsong, no dialogue.
Negative: No hands visible, no logos, no harsh reflections on glass.
4K 输出设置与宽高比
分辨率
Veo 3.1 支持多种输出分辨率。4K 目标为 3840×2160,即广播级 UHD 标准。更低分辨率(1080p、720p)也可使用,但无法发挥 Veo 3.1 在高端产品制作中的优势。
在 Vertex AI 中,需在生成参数中明确选择 4K;Gemini API 的默认分辨率可能较低。Google AI Studio Veo prompt cookbook 列出了各地区和账户类型支持的分辨率。
宽高比
| 宽高比 | 分辨率(4K) | 适用场景 |
|---|---|---|
| 16:9 | 3840×2160 | YouTube、广播、网页主横幅 |
| 9:16 | 2160×3840 | TikTok、Reels、Shorts、Stories |
| 1:1 | 2160×2160 | Instagram 信息流、PDP 画廊 |
| 4:5 | 2160×2700 | Instagram 信息流竖图、Pinterest |
| 21:9 | 3840×1620 | 电影信箱模式、高端网站主视觉 |
在提示词或 API 调用中指定宽高比。Veo 3.1 并不总是能从提示词中自动推断宽高比。
生成时间与成本
Veo 3.1 的 4K 生成比 1080p 更耗时且成本更高。建议先用 1080p 进行提示词迭代,最后在确认通过的提示词上以 4K 重新生成。这正是 Google AI Studio Veo prompt cookbook 中推荐的标准工作流。
规避 Veo 3.1 的薄弱环节
时长控制
Veo 3.1 的默认片段长度为 8 秒。更长的片段需要拼接生成结果,或在可用时使用模型的扩展片段模式。如果需要 15 秒的主镜头,建议生成两段 8 秒片段,并在第二段提示词中加入视觉连续性提示。
技巧:在第一段提示词结尾设置一个清晰的”结束姿态”(如手部停放、产品静止),以便第二段自然衔接。
多角色复杂度
两位画内说话者的同步对话是 Veo 3.1 中最容易失败的场景。模型可以视觉上生成两个角色,但唇形与特定台词的同步常常失配。变通方案:
- 一次只使用一个说话者,通过剪辑切换。
- 使用画外旁白代替画内对话。
- 仅使用模型生成画面,对话在后期添加。
文字与标志
Veo 3.1 仍然难以生成清晰的画面内文字。品牌标志经常呈现为模糊的近似形态。不要让 Veo 渲染你的文字商标,请在后期添加。
手部与手指瑕疵
与大多数当前视频模型一样,手部可能出现多余或合并的手指。负面提示”不要多余手指,不要变形的手”有助于改善,但并非总能生效。如果片段中手部是前景元素,建议安排人工审核。
多片段之间的一致性
对于需要在五个场景中展示同一产品的推广活动,建议先生成一段”主形象”片段,然后在后续提示词中描述该视觉结果(”与主形象镜头相同的哑光黑色瓶子,现在置于木质架子上”)。这种回溯参考技术在 Veo 提示词手册 中有详细记录。
常见问题
Veo 3.1 的最大片段长度是多少? Veo 3.1 默认生成 8 秒片段。扩展模式可根据你的 Vertex AI 或 Gemini API 等级生成更长片段。
Veo 3.1 真的能原生生成音频吗? 是的。对话、音效和环境声在视频推理的同一过程中生成,无需后期配音。
Veo 3.1 能输出 4K 吗? 可以,在支持的宽高比下最高可达 3840×2160。你必须在生成参数中明确选择 4K。
Veo 3.1 支持哪些宽高比? 16:9、9:16、1:1、4:5 和 21:9 为常见支持的比例。具体请根据你的 Vertex AI 地区和账户配置核实。
在产品视频方面,Veo 3.1 与 Sora 2 相比如何? 2025 年 11 月的 model comparison: Seedance vs Kling vs Veo 指出 Veo 3.1 的原生音频保真度更强。Sora 2 的默认片段往往更长,但同步对话效果较弱。
我可以将 Veo 3.1 的输出用于商业用途吗? 可以,依据 Google 生成式 AI 的标准条款。请务必在你签署的 Vertex AI 协议中核实最新的商业使用政策。
在哪里可以找到更多产品视频提示词示例? Google AI Studio Veo prompt cookbook 维护着一个包含 25 条提示词的产品库,Google AI Studio Veo prompt cookbook 则整理了面向商业用途的提示词。
如何让一个产品在不同片段之间保持一致? 使用回溯参考技术:先生成一段主形象镜头,然后在后续提示词中描述该视觉结果。Veo 提示词手册 对此有详细阐述。
结语
Veo 3.1 将原生音频和 4K 输出相结合,使其成为 2025 年末最强大的单提示词产品视频模型。该模型对摄影师级词汇、结构化提示词和明确音频分层均有良好响应。以上模板可作为起点——替换你的产品、调色板和音效,先在 1080p 下迭代,最后再提交 4K 生成。
如需了解更广泛的 AI 视频领域相关内容,请参阅我们关于 Seedance 2.5 角色动作提示词、Kling 3.0 角色一致性工作流 和 Wan 3.0 商业广告流水线 的指南。如需打下更扎实的提示词工程基础,2026 年最佳 AI 视频提示词大师课 和 2K 视频模型模板库 值得收藏。
官方 Google DeepMind Veo 模型页面 和 Google AI Studio Veo prompt cookbook 仍然是权威参考——请将它们加入书签,并在模型迭代时回顾查看。
由 videosprompt.org 编辑团队审校 · 2026 年 10 月
分享文章