用视频指导生成360度语音的空间音频,提升沉浸感。
Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes

- 通过视频定位声源位置,生成缺失的定向音频信号
- 在真实场景中实现更高精度的空间音频重建
- 适合做沉浸式媒体、虚拟现实的音频开发人员
空间音频是沉浸式360°媒体的关键组成部分,但在以语音为主的真实场景中,高质量的空间音频采集仍受限。本文研究了在真实环境中基于视觉引导的第一阶全向声(FOA)语音空间化:给定对齐的360°视频与全向音频轨道,恢复缺失的定向FOA分量。为此,我们从YouTube构建了面向语音的360°视频-FOA数据集YT-SPEECH。提出两阶段定位-渲染框架:基于音视频分割主干网络生成逐帧空间热图,再通过条件复域U-Net从全向通道重构定向FOA信号。采用置信度门控策略,在模糊声学条件下稳定条件输入。实验表明,相比消融版本和已有方法,本方法在重建保真度、空间准确性及感知语音质量上均有提升。
原文摘要 · Abstract (English)
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。