arXiv:2510.17788eess.AS2025-10被引 1

用音乐做声源,实现嘈杂环境下的鲁棒混响估计

AnyRIR: Robust Non-intrusive Room Impulse Response Estimation in the Wild

  • 以音乐为激励信号,时间-频率域上做稀疏约束的回归
  • 在真实噪声和编解码器不匹配下,性能优于传统方法
  • 适合AR/VR等需实时混响建模的应用场景

针对噪声干扰、非平稳声音(如语音或脚步)污染的开放环境中的房间冲激响应(RIR)估计问题,本文提出AnyRIR——一种非侵入式方法。该方法利用音乐作为激励信号替代专用测试信号,并在时频域将RIR估计建模为L1范数回归问题。通过迭代重加权最小二乘法(IRLS)与最小残差最小二乘法(LSMR)高效求解,充分利用非平稳噪声的稀疏性以抑制其影响。在仿真与实测数据上的实验表明,无论是在真实复杂场景还是编解码器不匹配条件下,AnyRIR均显著优于基于L2的时域与频域反卷积方法,为AR/VR等应用提供了鲁棒的RIR估计能力。

原文摘要 · Abstract (English)

We address the problem of estimating room impulse responses (RIRs) in noisy, uncontrolled environments where non-stationary sounds such as speech or footsteps corrupt conventional deconvolution. We propose AnyRIR, a non-intrusive method that uses music as the excitation signal instead of a dedicated test signal, and formulate RIR estimation as an L1-norm regression in the time-frequency domain. Solved efficiently with Iterative Reweighted Least Squares (IRLS) and Least-Squares Minimal Residual (LSMR) methods, this approach exploits the sparsity of non-stationary noise to suppress its influence. Experiments on simulated and measured data show that AnyRIR outperforms L2-based and frequency-domain deconvolution, under in-the-wild noisy scenarios and codec mismatch, enabling robust RIR estimation for AR/VR and related applications.

RIR估计音频处理稀疏恢复AR/VR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。