用视觉引导生成逼真空间音频,提升沉浸感。
SpatialV2A: Visual-Guided High-fidelity Spatial Audio Generation
- 构建首个大规模视频-双耳音频数据集,支持空间感知建模。
- 端到端框架显式建模空间特征,生成音频具真实空间深度。
- 在空间保真度上超越现有模型,适合沉浸式音视频应用。
尽管视频到音频生成在语义和时间对齐方面取得显著进展,但现有研究多局限于这些方面,对合成音频的空间感知与沉浸质量关注不足。这一局限主要源于当前模型依赖缺乏双耳空间信息的单声道音频数据集,难以学习视觉到空间音频的映射关系。为此,本文提出两项关键贡献:一是构建首个面向空间感知视频到音频生成的大规模视频-双耳音频数据集BinauralVGGSound;二是提出一种由视觉线索引导的端到端空间音频生成框架,显式建模空间特征。该框架包含视觉引导的音频空间化模块,确保生成音频具备真实的空间属性和层次化空间深度,同时保持语义与时间一致性。实验表明,本方法在空间保真度上显著优于现有先进模型,提供更沉浸的听觉体验,且不牺牲时间或语义一致性。演示页面见https://github.com/renlinjie868-web/SpatialV2A。
原文摘要 · Abstract (English)
While video-to-audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention to the spatial perception and immersive quality of the synthesized audio. This limitation stems largely from current models' reliance on mono audio datasets, which lack the binaural spatial information needed to learn visual-to-spatial audio mappings. To address this gap, we introduce two key contributions: we construct BinauralVGGSound, the first large-scale video-binaural audio dataset designed to support spatially aware video-to-audio generation; and we propose a end-to-end spatial audio generation framework guided by visual cues, which explicitly models spatial features. Our framework incorporates a visual-guided audio spatialization module that ensures the generated audio exhibits realistic spatial attributes and layered spatial depth while maintaining semantic and temporal alignment. Experiments show that our approach substantially outperforms state-of-the-art models in spatial fidelity and delivers a more immersive auditory experience, without sacrificing temporal or semantic consistency. The demo page can be accessed at https://github.com/renlinjie868-web/SpatialV2A.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。