从360度视频生成3D空间音频,还原声音方向感。
OmniAudio: Generating Spatial Audio from 360-Degree Video
- 用全景与视角视频双分支捕捉360视频的全局与局部信息。
- 在自监督预训练下,生成符合真实场景的空间音频(FOA格式)。
- 新数据集Sphere360支持高质量音视频对齐,适合沉浸式应用研究。
传统视频转音频方法多针对普通视角视频与非空间音频,缺乏三维环境中的声源方位信息。为解决此问题,本文提出新任务360V2SA,即从360度视频生成空间音频,输出标准的一阶全向声学(FOA)格式,可精确表达声音方向并实现真实3D音频还原。我们构建了名为Sphere360的新数据集,基于真实世界数据,并设计高效半自动音视频配对采集与清洗流程。为此,提出OmniAudio框架,采用自监督预训练,融合空间音频(FOA)与大规模非空间音频数据。该框架具备双分支结构,同时利用全景与视角视频输入,全面捕捉360视频的时空特征。实验表明,OmniAudio在Sphere360数据集上于客观与主观评价指标均达当前最优表现。代码与数据集已开源。
原文摘要 · Abstract (English)
Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。