arXiv:2607.01834cs.SD2026-07中稿 · Interspeech 2026被引 1

低功耗助听器实现实时双耳语音增强,延迟低至8毫秒。

RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices

论文配图:RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices
图 1 · 摘自论文原文
  • 分阶段分布式架构,结合感知特征压缩与轻量递归掩码估计
  • 计算量大幅降低,8毫秒超低延迟下仍保持良好音质
  • 适合资源受限的可穿戴设备,尤其适用于实时助听场景

实时双耳语音增强受限于延迟、计算开销和设备间通信,现有高效方案多集中于单通道场景。本文提出RT-Tango,一种专为资源受限平台(如助听器)设计的实时分布式双耳语音增强框架。该框架采用两阶段分布式架构,结合感知驱动的ERB特征压缩、轻量级分组递归掩码估计与时间稀疏化,显著降低计算成本。通过非对称STFT解耦谱分辨率与算法延迟,配合因果递归推理与空间统计在线估计,有效满足严格延迟约束。实验表明,RT-Tango在保持优异语音增强性能的同时,大幅减少MACs运算量,并实现最低达8毫秒的超低延迟。

原文摘要 · Abstract (English)

Real-time binaural speech enhancement is constrained by latency, computational cost, and inter-device communication, yet existing efficient solutions predominantly address single-channel settings. In this paper, we introduce RT-Tango, a real-time distributed binaural speech enhancement framework designed for streaming on resource-constrained platforms and specifically for hearing aids. RT-Tango relies on a two-stage distributed architecture combining perceptually motivated ERB feature compression, lightweight grouped recurrent mask estimation, and temporal sparsification to reduce computational cost. Stringent latency constraints are addressed by decoupling spectral resolution from algorithmic delay using an asymmetric STFT, together with causal recurrent inference and online estimation of spatial statistics. Experimental results show that RT-Tango achieves competitive speech enhancement while significantly reducing MACs operations and functioning at ultra-low latencies as low as 8 ms.

语音增强双耳处理低延迟助听器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。