arXiv:2510.20602cs.SDcs.AI2025-10NeurIPS被引 2

用声学互易性实现任意位置发声的沉浸式音效建模

Resounding Acoustic Fields with Reciprocity

  • 通过互换声源与听者位置,生成密集虚拟声源的物理合理样本
  • 在模拟与真实数据上均显著提升声场学习性能,感知实验验证沉浸感增强
  • 解决互易性部署中增益不匹配问题,自监督学习提升鲁棒性

在虚拟环境中实现沉浸式听觉体验需要支持动态声源位置的灵活声音建模。本文提出一项名为「resounding」的任务,旨在从稀疏测量的声源位置估计任意位置的房间冲击响应,类似于视觉中的再照明问题。我们利用声学互易性,提出Versa——一种受物理启发的方法,通过交换声源与听者姿态,生成具有密集虚拟声源位置的物理有效样本。同时,我们识别出互易性部署中的声源/听者增益模式挑战,并提出自监督学习方法加以解决。实验结果表明,Versa在模拟和真实数据集上均显著提升声场学习性能,多种指标表现优异。感知用户研究表明,Versa可大幅改善空间音频沉浸感。代码、数据集及演示视频已发布于项目网站:https://waves.seas.upenn.edu/projects/versa。

原文摘要 · Abstract (English)

Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called resounding, which aims to estimate room impulse responses at arbitrary emitter location from a sparse set of measured emitter positions, analogous to the relighting problem in vision. We leverage the reciprocity property and introduce Versa, a physics-inspired approach to facilitating acoustic field learning. Our method creates physically valid samples with dense virtual emitter positions by exchanging emitter and listener poses. We also identify challenges in deploying reciprocity due to emitter/listener gain patterns and propose a self-supervised learning approach to address them. Results show that Versa substantially improve the performance of acoustic field learning on both simulated and real-world datasets across different metrics. Perceptual user studies show that Versa can greatly improve the immersive spatial sound experience. Code, dataset and demo videos are available on the project website: https://waves.seas.upenn.edu/projects/versa.

声场建模互易性空间音频自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。