arXiv:2502.07538cs.MMcs.SD2025-02被引 3

用视觉信息自动生成多说话人场景的空间音频,省去人工对齐。

Visual-based spatial audio generation system for multi-speaker environments

  • 基于YOLOv8和单目深度估计,从视频推断说话人位置。
  • 无需额外训练数据,音频与视频空间一致性显著提升。
  • 适合影视游戏制作中快速生成高质量空间音频。

在电影和视频游戏等多媒体应用中,空间音频技术广泛用于增强用户体验,将单声道音频转换为双耳格式以模拟三维声音效果。然而,这一过程通常复杂且耗时,需声效设计师精确同步音频与视觉元素的空间位置。为此,我们提出一种基于视觉的空间音频生成系统——该系统集成YOLOv8人脸检测、单目深度估计与空间音频技术,无需额外双耳数据集训练。通过客观指标评估,实验结果表明,本方法显著提升了音频与视频间的空间一致性,改善了语音质量,并在多说话人场景下表现稳健。该系统简化了音画对齐流程,使声效工程师能高效生成高质量音频,是多媒体制作领域的重要工具。

原文摘要 · Abstract (English)

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex and labor-intensive for sound designers, requiring precise synchronization of audio with the spatial positions of visual components. To address these challenges, we propose a visual-based spatial audio generation system - an automated system that integrates face detection YOLOv8 for object detection, monocular depth estimation, and spatial audio techniques. Notably, the system operates without requiring additional binaural dataset training. The proposed system is evaluated against existing Spatial Audio generation system using objective metrics. Experimental results demonstrate that our method significantly improves spatial consistency between audio and video, enhances speech quality, and performs robustly in multi-speaker scenarios. By streamlining the audio-visual alignment process, the proposed system enables sound engineers to achieve high-quality results efficiently, making it a valuable tool for professionals in multimedia production.

空间音频视觉生成多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。