让虚拟声音随环境实时变化,提升沉浸感。
Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
- 融合房间结构、材质与语义信息,动态建模声学环境。
- 生成逼真的混响响应,音质在多场景中表现稳定。
- 适合做虚拟现实/增强现实音频优化的研究者与开发者。
在扩展现实(XR)中,还原真实世界的声学效果对构建身临其境的虚拟体验至关重要。然而,现有空间音频渲染方法难以实时适配多样物理场景,导致视听感知不一致,破坏沉浸感。为此,我们提出SAMOSA——一种基于设备端的新型系统,通过融合实时估计的房间几何、表面材料及语义驱动的声学上下文,构建协同的多模态场景表示。该表示利用场景先验实现高效声学校准,从而合成高度真实的混响脉冲响应(RIR)。我们在多种房间配置和声源类型下,通过声学指标评估了RIR合成性能,并开展专家评测(N=12)。结果表明,SAMOSA在提升XR听觉真实感方面具备可行性与有效性。
原文摘要 · Abstract (English)
In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。