arXiv:2608.25596eess.AS2026-08中稿 · EUSIPCO 2026

用知识蒸馏让语音回声消除模型变小100倍,性能还更好。

Knowledge Distillation for Efficient Acoustic Echo Control

  • 用知识蒸馏压缩复杂模型,保留核心性能。
  • 仅2%计算量下近端语音失真显著减少。
  • 比六倍大的模型更优,适合移动端部署。

近年来,机器学习方法逐渐取代传统自适应滤波器用于声学回声消除(AEC)。尽管性能超越经典算法成为可能,但计算复杂度仍是主要挑战。典型架构如卷积循环网络(CRNs)的计算开销是传统信号处理方案的数个数量级。虽然模型压缩易于实现,但常导致性能明显下降。本文首次在AEC领域证明,通过有效知识蒸馏(KD)可大幅缓解性能损失,实现高性能高效AEC。所提CGGN16学生模型在仅2%教师模型计算复杂度下,近端语音失真显著降低,整体性能超过使用真实标签训练的六倍复杂度模型,并优于近期文献中的其他专用AEC架构。

原文摘要 · Abstract (English)

In recent years, many efforts have been made to supersede classical acoustic echo control (AEC) algorithms with more powerful machine-learned approaches. While surpassing the performance of well-established adaptive filters is very much possible, a remaining challenge is computational complexity. Popular architectures, such as convolutional recurrent networks (CRNs), are by multiple orders of magnitude computationally more expensive than classical signal processing solutions. Scaling down such models is usually straight-forward, but it comes at the cost of a notably reduced performance. We show - to the author's knowledge for the first time in AEC - how these performance drops can be successfully alleviated to a large degree by employing an effective knowledge distillation (KD) process, enabling more potent efficient AEC. Our proposed CGGN16 student AEC models show significantly less near-end speech distortion at only 2% of its teacher's computational complexity, surpass the overall performance of a six times more complex model trained on ground-truth labels, and outperform other AEC-focused architectures from recent literature.

回声消除知识蒸馏高效模型语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。