轻量级神经网络解决手机语音通话回声问题
A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
- 用数据增强和渐进学习提升模型抗干扰能力
- 在真实场景下显著改善语音质量与识别率
- 适合资源受限的移动端实时语音应用
在全双工语音交互系统中,有效的声学回声消除(AEC)对恢复受回声污染的语音至关重要。本文提出一种基于神经网络的AEC解决方案,以应对移动场景中硬件差异、非线性失真和长延迟带来的挑战。首先,通过多种数据增强策略提升模型在不同环境下的鲁棒性;其次,采用渐进式学习逐步提升AEC效果,显著改善语音质量。为优化下游应用,引入针对语音活动检测(VAD)和自动语音识别(ASR)任务定制的后处理策略,进一步提升其性能。最终方法采用小尺寸模型并支持流式推理,实现移动端无缝部署。实证结果表明,该方法在回声返回损耗增益(ERLE)和语音感知质量评价(PESQ)方面表现优异,并在VAD和ASR任务上均取得显著提升。
原文摘要 · Abstract (English)
In full-duplex speech interaction systems, effective Acoustic Echo Cancellation (AEC) is crucial for recovering echo-contaminated speech. This paper presents a neural network-based AEC solution to address challenges in mobile scenarios with varying hardware, nonlinear distortions and long latency. We first incorporate diverse data augmentation strategies to enhance the model's robustness across various environments. Moreover, progressive learning is employed to incrementally improve AEC effectiveness, resulting in a considerable improvement in speech quality. To further optimize AEC's downstream applications, we introduce a novel post-processing strategy employing tailored parameters designed specifically for tasks such as Voice Activity Detection (VAD) and Automatic Speech Recognition (ASR), thus enhancing their overall efficacy. Finally, our method employs a small-footprint model with streaming inference, enabling seamless deployment on mobile devices. Empirical results demonstrate effectiveness of the proposed method in Echo Return Loss Enhancement and Perceptual Evaluation of Speech Quality, alongside significant improvements in both VAD and ASR results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。