提出新型轻量级回声消除框架,提升压缩特征下的时频建模能力。
Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation

- 采用双分支编码器与回声感知调制模块增强压缩特征
- 仅增加26.1%计算量即达99.1% PESQ,超越频率域模型
- 适合资源受限设备,参数仅0.2M,每秒100M FLOPs
现有轻量级回声消除系统常结合线性回声消除与巴克域深度神经网络抑制以降低计算开销。此类系统通过下采样层将输入特征压缩为紧凑瓶颈表示,但该压缩削弱了时频建模能力并导致性能下降。为此,本文提出MSA-EchoLite:一种轻量级巴克域回声消除框架,包含非对称双分支编码器与回声感知时频调制(EAM)模块。EAM模块通过建模双分支麦克风与回声相关潜在特征间的差异与关联线索,丰富压缩后的瓶颈表示。实验表明,相较于频率域版本,巴克域变体在性能-复杂度权衡上更优,但对特征压缩更敏感。其增强版仅比无EAM的巴克域版本多出26.1%浮点运算量,却达到频率域版本99.1%的PESQ分数(后者需近两倍浮点运算量),且在SDR上表现更优。整体上,MSA-EchoLite优于现有轻量级回声消除模型,仅使用0.2M参数与100M FLOPs/s。
原文摘要 · Abstract (English)
Existing lightweight acoustic echo cancellation (AEC) systems often combine linear AEC with Bark-domain DNN-based suppression to lower the computational footprint. In such systems, downsampling layers further compress the input features into a compact bottleneck representation, but this compression weakens frequency-time modeling capacity and degrades performance. To mitigate this limitation, we propose MSA-EchoLite, a lightweight Bark-domain AEC framework with an asymmetric dual-branch encoder and an echo-aware frequency-time modulation (EAM) module. The EAM module enriches the compressed bottleneck representation by modeling discrepancy and correlation cues between the dual-branch microphone and echo-related latent features. Experimental results show that the Bark-domain variant of MSA-EchoLite offers a better performance-complexity trade-off than its frequency-domain counterpart but is more sensitive to feature compression. With only 26.1% additional FLOPs over its non-EAM Bark-domain variant, its EAM-enhanced version achieves 99.1% of the PESQ of the frequency-domain counterpart, which requires nearly twice the FLOPs, and even surpasses it in SDR. Overall, MSA-EchoLite outperforms state-of-the-art lightweight AEC models while using only 0.2 M parameters and 100 M FLOPs/s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。