arXiv:2409.18705eess.AScs.AI2024-09中稿 · Interspeech 2024被引 2

为真无线耳机设计低延迟语音增强,提升嘈杂环境通话质量。

Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds

  • 采用轻量网络架构与硬件优化,降低计算开销。
  • 算法延迟低于3毫秒,支持实时对话。
  • 适合需要清晰语音的耳机用户和嵌入式设备开发者。

本文提出一种专为真无线立体声(TWS)耳机设计的本地化语音增强方案,旨在支持开启主动降噪(ANC)时的嘈杂环境通话。主要挑战在于模型计算复杂度高且算法延迟必须低于3毫秒,以维持实时对话体验。为此,我们评估了网络架构、损失函数设计、剪枝方法及硬件特定优化等多个关键要素。实验表明,该方案在保持低延迟的同时,显著优于基线模型的语音增强效果,同时大幅降低计算复杂度。

原文摘要 · Abstract (English)

This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations in noisy environments, with active noise cancellation (ANC) activated. The primary challenges for speech enhancement models in this context arise from computational complexity that limits on-device usage and latency that must be less than 3 ms to preserve a live conversation. To address these issues, we evaluated several crucial design elements, including the network architecture and domain, design of loss functions, pruning method, and hardware-specific optimization. Consequently, we demonstrated substantial improvements in speech enhancement quality compared with that in baseline models, while simultaneously reducing the computational complexity and algorithmic latency.

语音增强低延迟TWS耳机嵌入式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。