arXiv:2508.12918cs.SD2025-08被引 3

用视觉信息生成沉浸式立体声,让声音随画面动起来

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

  • 根据视频帧定位声源2D位置与深度,映射为3D轨迹
  • 结合预训练模型的单声道音频,生成空间一致的立体声
  • 自建含动态声源的头相关脉冲响应数据集,提升真实感

随着AIGC的发展,基于深度学习的视频到音频(V2A)技术受到广泛关注。然而,现有研究多聚焦于缺乏空间感知的单声道音频生成,对能提供更强沉浸感的双耳空间音频生成探索不足。为此,我们提出FoleySpace框架,利用视觉信息生成沉浸且空间一致的立体声。具体而言,我们设计了一种声源估计方法,确定每帧视频中声源的2D坐标和深度,并通过坐标映射机制将其转换为3D轨迹。该3D轨迹与预训练V2A模型生成的单声道音频共同作为条件输入,驱动扩散模型生成空间一致的双耳音频。为支持动态声场生成,我们基于实测头相关脉冲响应(Head-Related Impulse Responses)构建了包含多种声源运动场景的训练数据集。实验表明,所提方法在空间感知一致性上优于现有方法,显著提升了音画体验的沉浸感。

原文摘要 · Abstract (English)

Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono audio generation that lacks spatial perception, while the exploration of binaural spatial audio generation technologies, which can provide a stronger sense of immersion, remains insufficient. To solve this problem, we propose FoleySpace, a framework for video-to-binaural audio generation that produces immersive and spatially consistent stereo sound guided by visual information. Specifically, we develop a sound source estimation method to determine the sound source 2D coordinates and depth in each video frame, and then employ a coordinate mapping mechanism to convert the 2D source positions into a 3D trajectory. This 3D trajectory, together with the monaural audio generated by a pre-trained V2A model, serves as a conditioning input for a diffusion model to generate spatially consistent binaural audio. To support the generation of dynamic sound fields, we constructed a training dataset based on recorded Head-Related Impulse Responses that includes various sound source movement scenarios. Experimental results demonstrate that the proposed method outperforms existing approaches in spatial perception consistency, effectively enhancing the immersive quality of the audio-visual experience.

音频生成空间音频视觉引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。