轻量级语音回声消除模型,参数仅278K,性能媲美顶尖方案。
EchoFree: Towards Ultra Lightweight and Efficient Neural Acoustic Echo Cancellation
- 用线性滤波加巴克频带神经后处理,降低计算负担。
- 仅278K参数、30MMACs算力,仍超越现有轻量模型。
- 适合移动端实时通话等低延迟场景使用。
近年来,神经网络被广泛应用于声学回声消除(AEC)。然而,现有方法难以在保持性能的同时满足真实场景下的低延迟与低算力要求。为此,我们提出EchoFree,一种超轻量级神经AEC框架,结合线性滤波与神经后处理。具体地,设计了一个在巴克频带谱特征上运行的神经后滤波器,并引入两阶段优化策略,利用自监督学习(SSL)模型提升性能。我们在ICASSP 2023 AEC挑战赛的盲测集上评估该方法。结果表明,该模型仅需278K参数和30MMACs计算复杂度,优于现有轻量级AEC模型,且性能接近当前最先进的轻量级模型DeepVQE-S。音频示例已公开。
原文摘要 · Abstract (English)
In recent years, neural networks (NNs) have been widely applied in acoustic echo cancellation (AEC). However, existing approaches struggle to meet real-world low-latency and computational requirements while maintaining performance. To address this challenge, we propose EchoFree, an ultra lightweight neural AEC framework that combines linear filtering with a neural post filter. Specifically, we design a neural post-filter operating on Bark-scale spectral features. Furthermore, we introduce a two-stage optimization strategy utilizing self-supervised learning (SSL) models to improve model performance. We evaluate our method on the blind test set of the ICASSP 2023 AEC Challenge. The results demonstrate that our model, with only 278K parameters and 30 MMACs computational complexity, outperforms existing low-complexity AEC models and achieves performance comparable to that of state-of-the-art lightweight model DeepVQE-S. The audio examples are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。