跨环境实时还原空间声学,仅需少量测量
Hearing Anywhere in Any Environment
- 结合全景深度图与少量参考声学数据建模
- 在30万+真实感模拟场景中表现远超基线
- 适合混合现实、声学仿真等需要泛化能力的场景
在混合现实应用中,空间环境中的真实听觉体验与视觉沉浸同样关键。尽管神经网络在房间冲激响应(RIR)估计方面取得进展,但多数方法仅限于训练过的单一环境,无法推广到几何结构和表面材料不同的新房间。本文提出xRIR框架,实现跨房间RIR预测,构建可泛化至任意环境的统一模型,仅需少量额外测量即可重建空间声学。核心思想是融合几何特征提取器(从全景深度图捕获空间上下文)与RIR编码器(从少数参考RIR样本中提取精细声学特征)。为评估方法,我们构建ACOUSTICROOMS数据集,包含260个房间生成的超过30万组高保真模拟RIR。实验表明,该方法显著优于多个基线。此外,通过在4个真实环境上进行模拟到真实迁移测试,验证了方法的泛化能力及数据集的真实性。
原文摘要 · Abstract (English)
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present xRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce ACOUSTICROOMS, a new dataset featuring high-fidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four real-world environments, demonstrating the generalizability of our approach and the realism of our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。