arXiv:2509.15492cs.SDcs.MM2025-09被引 1

让视频生成带环境感的清晰语音,同步性强且自然。

Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech

  • 分两阶段建模:先预测含语音线索的语义音符,再转为具体音频。
  • 在动态视频上实现语音与画面精准对齐,支持沉浸式背景音转换。
  • 适合游戏、影视等需高同步性语音合成的场景,尤其关注环境感知。

真实世界应用如游戏开发中,生成逼真且上下文感知的音频至关重要。现有视频到音频(V2A)方法多聚焦于拟音生成,难以产出可理解的语音;而当前环境语音合成仍依赖文本输入,无法与动态视频内容时序对齐。本文提出超越视频到拟音(BVS)的方法,可为给定视频生成同步且具备环境感知的清晰语音。采用两阶段建模:第一阶段为视频引导的音频语义模型(V2AS),基于语音线索预测统一的音频语义标记;第二阶段为视频条件的语义到声学模型(VS2A),将语义标记细化为详细声学标记。实验表明,BVS在视频到上下文感知语音合成及沉浸式背景音转换等场景中表现优异,消融实验进一步验证了设计的有效性。演示链接:https://xinleiniu.github.io/BVS-demo/

原文摘要 · Abstract (English)

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce intelligible speech. Meanwhile, current environmental speech synthesis approaches remain text-driven and fail to temporally align with dynamic video content. In this paper, we propose Beyond Video-to-SFX (BVS), a method to generate synchronized audio with environmentally aware intelligible speech for given videos. We introduce a two-stage modeling method: (1) stage one is a video-guided audio semantic (V2AS) model to predict unified audio semantic tokens conditioned on phonetic cues; (2) stage two is a video-conditioned semantic-to-acoustic (VS2A) model that refines semantic tokens into detailed acoustic tokens. Experiments demonstrate the effectiveness of BVS in scenarios such as video-to-context-aware speech synthesis and immersive audio background conversion, with ablation studies further validating our design. Our demonstration is available at~\href{https://xinleiniu.github.io/BVS-demo/}{BVS-Demo}.

语音合成视频生成环境感知同步音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。