融合物理模型与统计方法,提升混响语音的恢复效果。
Modèle physique variationnel pour l'estimation de réponses impulsionnelles de salles
- 将混响响应分解为可解释的物理参数,结合统计建模
- 在噪声环境下优于传统去卷积方法,客观指标更优
- 适合语音增强与声学建模研究者参考
房间脉冲响应(RIR)估计对语音去混响等任务至关重要,有助于提升自动语音识别性能。现有方法多依赖统计信号处理或模仿信号处理原理的深度神经网络,但将统计与物理建模结合用于RIR估计的研究仍较少。本文提出一种新方法,通过理论严谨的模型整合两者:将RIR分解为白高斯噪声经频率相关指数衰减(模拟墙面吸音)及自回归滤波器(模拟麦克风响应)的组合;利用变分自由能代价函数实现参数估计。以干声与混响语音为输入,验证表明该方法在噪声环境中优于经典去卷积,客观指标表现更佳。
原文摘要 · Abstract (English)
Room impulse response estimation is essential for tasks like speech dereverberation, which improves automatic speech recognition. Most existing methods rely on either statistical signal processing or deep neural networks designed to replicate signal processing principles. However, combining statistical and physical modeling for RIR estimation remains largely unexplored. This paper proposes a novel approach integrating both aspects through a theoretically grounded model. The RIR is decomposed into interpretable parameters: white Gaussian noise filtered by a frequency-dependent exponential decay (e.g. modeling wall absorption) and an autoregressive filter (e.g. modeling microphone response). A variational free-energy cost function enables practical parameter estimation. As a proof of concept, we show that given dry and reverberant speech signals, the proposed method outperforms classical deconvolution in noisy environments, as validated by objective metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。