用神经网络建模车载空间声学响应,自动捕捉频率特性与动态变化。
INFER : Learning Implicit Neural Frequency Response Fields for Confined Car Cabin
- 基于频域的隐式神经场,联合输入声源与接收器位置方向,学习三维复数声响应。
- 在真实汽车舱内数据上,幅值和相位重建误差分别降低39%和51%。
- 融合听觉感知与物理规律,适合车载音频系统实时优化与个性化调校。
精确建模封闭共振环境(如汽车舱内)的空间声学对实现沉浸式、可懂性强的音频至关重要。现有调音方法依赖人工、硬件密集且静态,无法反映频率选择性行为及乘客存在、座椅调整等动态变化。为此,我们提出INFER:隐式神经频率响应场,一种联合条件于声源与接收器位置、朝向的频域神经框架,可直接学习汽车舱等封闭共振环境内的复数频率响应场。相比现有神经声学建模方法,提出三项关键创新:(1) 新型端到端频域前向模型,直接学习三维空间中的频率响应场与频率相关衰减;(2) 感知与硬件感知的谱监督机制,强调关键听觉频段,弱化不稳定的跨带区域;(3) 基于克雷默-克龙尼(Kramers-Kronig)一致性的物理约束,正则化频率相关的衰减与延迟。我们在多个真实汽车舱内采集的数据上评估该方法,显著优于时域与混合域基线,在模拟与真实汽车数据集上平均幅值和相位重建误差分别降低超过39%和51%。INFER为车载空间神经声学建模设立了新基准。
原文摘要 · Abstract (English)
Accurate modeling of spatial acoustics is critical for immersive and intelligible audio in confined, resonant environments such as car cabins. Current tuning methods are manual, hardware-intensive, and static, failing to account for frequency selective behaviors and dynamic changes like passenger presence or seat adjustments. To address this issue, we propose INFER: Implicit Neural Frequency Response fields, a frequency-domain neural framework that is jointly conditioned on source and receiver positions, orientations to directly learn complex-valued frequency response fields inside confined, resonant environments like car cabins. We introduce three key innovations over current neural acoustic modeling methods: (1) novel end-to-end frequency-domain forward model that directly learns the frequency response field and frequency-specific attenuation in 3D space; (2) perceptual and hardware-aware spectral supervision that emphasizes critical auditory frequency bands and deemphasizes unstable crossover regions; and (3) a physics-based Kramers-Kronig consistency constraint that regularizes frequency-dependent attenuation and delay. We evaluate our method over real-world data collected in multiple car cabins. Our approach significantly outperforms time- and hybrid-domain baselines on both simulated and real-world automotive datasets, cutting average magnitude and phase reconstruction errors by over 39% and 51%, respectively. INFER sets a new state-of-the-art for neural acoustic modeling in automotive spaces
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。