用神经网络改进声学回声消除中的卡尔曼滤波,提升收敛速度与抗干扰能力。
Neural Kalman Filters for Acoustic Echo Cancellation
- 将深度神经网络嵌入频域卡尔曼滤波,自动估计噪声协方差
- 在双讲场景下保留近端语音效果优于传统方法
- 适用于对实时性与语音质量要求高的免提通信系统
卡尔曼滤波是信号处理中自适应滤波的强大工具。频域自适应卡尔曼滤波(FDKF)基于声学状态空间模型,统一了自适应滤波更新与步长控制,专为声学回声消除设计,广泛应用于免提系统。本文回顾线性FDKF,并探讨如何通过深度神经网络(DNN)增强其性能,特别是克服传统方法中需手动估计过程与观测噪声协方差的难题。尽管纯FDKF计算开销极低,但其神经卡尔曼滤波变体可实现更快的(重)收敛、更优的回声消除效果,并在在线性与非线性扬声器条件下均优于原方法,尤其在双讲时仍能良好保留近端语音。本文在同一训练框架与数据集上对比多种DNN增强型FDKF方案,提供当前技术前沿的综合分析。
原文摘要 · Abstract (English)
Kalman filtering is a powerful approach to adaptive filtering for various problems in signal processing. The frequency-domain adaptive Kalman filter (FDKF), based on the concept of the acoustic state space, provides a unifying solution to the adaptive filter update and the related stepsize control. It was conceived for the problem of acoustic echo cancellation and, as such, is frequently applied in hands-free systems. This article motivates and briefly recapitulates the linear FDKF and investigates how it can be further supported by deep neural networks (DNNs) in various ways, specifically to overcome the challenges and limitations related to the usually required estimation of process and observation noise covariances for the Kalman filter. While the mere FDKF comes with very low computational complexity, its neural Kalman filter variants may deliver faster (re)convergence, better echo cancellation, and even exceed the FDKF in its excellent double-talk near-end speech preservation both under linear and nonlinear loudspeaker conditions. To provide a synopsis of the state of the art, this article contributes a comparison of a range of DNN-based extensions of FDKF in the same training framework and using the same data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。