arXiv:2605.20968eess.ASeess.SP2026-05

用神经网络从房间结构预测声音衰减曲线,更真实还原听感。

From Numbers to Perception, Energy Decay Curves Prediction

论文配图:From Numbers to Perception, Energy Decay Curves Prediction
图 1 · 摘自论文原文
  • 直接从房间几何和材料预测多频段能量衰减曲线。
  • 自定义复合损失函数使预测结果符合物理衰减规律。
  • 适合虚拟现实音频渲染,计算效率高且听感准确。

由于音频信号维度高且需感知准确性,预测混响冲击响应(RIR)仍具挑战。本文提出一种神经网络框架,直接从房间几何与材料属性预测多频段能量衰减曲线(EDCs)。不同于传统模型,该框架采用自定义复合损失函数,在对数域中同时优化能量水平与衰减斜率,确保预测结果符合物理衰减规律,并保持对混响时间T30和早期反射的高敏感性。实验表明,模型在T30和清晰度指数上误差极小,能有效逼近真实声学特性。该方法为传统仿真提供了计算高效的替代方案,适用于交互式虚拟环境中的逼真音频渲染。

原文摘要 · Abstract (English)

Predicting Room Impulse Responses (RIRs) remains a challenge due to the high dimensionality of audio signals and the need for perceptual accuracy. This paper introduces a neural network framework that predicts multi-band Energy Decay Curves (EDCs) directly from room geometry and material properties. Unlike standard models, our framework employs a custom composite loss function that optimizes for both energy levels and decay slopes in the log-domain. This ensures the predicted curves adhere to physical decay principles while maintaining high sensitivity to reverberation time and early reflections. Results demonstrate that the model successfully approximates ground-truth acoustics with minimal error in T30 and clarity indices. The approach offers a computationally efficient alternative to traditional simulations, facilitating realistic audio rendering for interactive virtual environments.

声学建模神经网络虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。