arXiv:2607.02062eess.AS2026-07中稿 · Interspeech 2026

轻量级网络联合消除回声与降噪,适配设备端实时语音对话。

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

论文配图:LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression
图 1 · 摘自论文原文
  • 多路径对齐修正参考信号与麦克风信号的时间和能量偏差。
  • 动态注意力融合增强的线性回声消除与麦克风特征,适应不同环境。
  • 仅480K参数、126MACs,实现实时推理且性能媲美顶尖轻量模型。

我们提出一种轻量级多路径对齐网络(LMPAN),用于全双工语音对话系统中的设备端联合回声消除(AEC)与噪声抑制(NS)。为应对硬件引起的失真和动态声学环境,提出三项核心创新:(1)多路径对齐阶段校正参考信号、线性回声消除(LAEC)输出与麦克风信号间的时间与能量差异;(2)基于注意力机制,在不同声学场景下动态融合增强的LAEC与麦克风特征;(3)采用动态目标生成策略的后处理模块,提升下游任务(如自动语音识别、语音活动检测)性能。此外,采用两阶段训练框架,利用自监督学习表示以提升感知质量。实验表明,LMPAN仅需480K参数和126 MACs,性能可媲美当前最先进的轻量模型DeepVQE-S,同时保证实时推理能力。

原文摘要 · Abstract (English)

We propose a lightweight multi-path alignment network (LMPAN) for on-device joint acoustic echo cancellation (AEC) and noise suppression (NS) in full-duplex spoken dialogue systems. To address hardware-induced distortions and dynamic acoustic conditions, we introduce three core innovations: (1) a multi-path alignment stage correcting temporal and energy mismatches across reference, linear AEC (LAEC) output, and microphone signals; (2) an attention-based mechanism that dynamically integrates enhanced LAEC and microphone features under varying acoustic scenarios; (3) a post-filtering module with a dynamic target generation strategy for downstream tasks (ASR, VAD). Furthermore, we adopt a two-stage training framework leveraging self-supervised learning representations to enhance perceptual quality. Experiments show that LMPAN, with only 480K parameters and 126 MACs, achieves performance comparable to the state-of-the-art lightweight model DeepVQE-S, while ensuring real-time inference capability.

语音处理轻量模型回声消除降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。