用地图和少量测量数据,快速估算虚拟场景中各位置的声学参数。
Scene-wide Acoustic Parameter Estimation
- 通过2D平面图+校准声学数据,生成全场景声学热力图。
- 在1000个复杂房间数据集上,性能优于传统统计模型。
- 支持定向声学参数预测,适合沉浸式音视频应用。
在增强现实(AR)和虚拟现实(VR)应用中,准确估计场景声学特性对提升沉浸感至关重要。然而,仅从场景几何结构直接推导混响脉冲响应(RIR)通常计算复杂且数据成本高。本文提出一种新方法,利用AR/VR环境中易获取的轻量信息,推断整个场景的空间分布声学参数(如C50、T60等)。我们将问题建模为图像到图像的转换任务,将带校准RIR的2D平面图转化为声学参数的2D热力图。此外,我们证明该方法同样适用于方向相关(即波束成形)参数的预测。为此,我们构建并发布了包含1000个复杂场景的公开数据集,用于研究该任务,并验证其在性能上优于强统计基线。
原文摘要 · Abstract (English)
For augmented (AR) and virtual reality (VR) applications, accurate estimates of the acoustic characteristics of a scene are critical for creating a sense of immersion. However, directly estimating Room-impulse Responses (RIRs) from scene geometry is often a challenging, data-expensive task. We propose a method to instead infer spatially-distributed acoustic parameters (such as C50, T60, etc) for an entire scene from lightweight information readily available in an AR/VR context. We consider an image-to-image translation task to transform a 2D floormap, conditioned on a calibration RIR measurement, into 2D heatmaps of acoustic parameters. Moreover, we show that the method also works for directionally-dependent (i.e. beamformed) parameter prediction. We introduce and release a 1000-room, complex-scene dataset to study the task, and demonstrate improvements over strong statistical baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。