arXiv:2509.15210cs.SDcs.AI2025-09

用房间几何结构指导神经声学建模,生成更逼真的混响响应。

Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation

  • 基于粗略房间网格查询局部距离分布,显式编码几何上下文。
  • 在多个评估指标上优于传统与前沿方法,提升混响预测精度。
  • 适合需要高保真声学模拟的虚拟现实与音频渲染场景。

真实感声音模拟在诸多应用中至关重要,其中房间冲激响应(RIR)是表征声音在特定空间中传播特性的核心要素。近期研究利用神经隐式方法,通过环境中的场景图像等上下文信息学习RIR,但未能有效利用环境中的显式几何信息。为更好地结合神经隐式模型与直接几何特征,本文提出MiNAF:在给定位置查询粗略房间网格,并提取距离分布作为局部上下文的显式表示。实验表明,引入显式局部几何特征能更有效地引导模型生成更准确的RIR预测。与传统及最先进方法相比,MiNAF在多种评估指标上表现具有竞争力。

原文摘要 · Abstract (English)

Realistic sound simulation plays a critical role in many applications. A key element in sound simulation is the room impulse response (RIR), which characterizes how sound propagates within a given space. Recent studies have applied neural implicit methods to learn RIR using context information collected from the environment, such as scene images. However, these approaches do not effectively leverage explicit geometric information from the environment. To further exploit neural implicit models with direct geometric features, we present MiNAF, which queries a rough room mesh at given locations and extracts distance distributions as an explicit representation of local context. Our approach demonstrates that incorporating explicit local geometric features can better guide the model in generating more accurate RIR predictions. Through comparisons with conventional and state-of-the-art methods, we show that MiNAF performs competitively across various evaluation metrics.

声学建模神经隐式几何特征混响生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。