arXiv:2508.03047cs.SDcs.LG2025-08被引 4

首个能在低功耗耳机芯片上实时运行的语音分离模型。

TF-MLPNet: Tiny Real-Time Neural Speech Separation

  • 在时频域用全连接层交替处理通道与频率,结合卷积处理时间序列。
  • 6毫秒音频块实时处理,相比之前模型提速3.5至4倍。
  • 适合资源受限的可穿戴设备,尤其适用于助听场景。

可穿戴设备上的语音分离技术可实现变革性的增强听力功能。然而,现有先进语音分离网络因计算量大,无法在为可穿戴设备设计的低功耗神经加速器上实时运行。本文提出TF-MLPNet,首个可在该类低功耗加速器上实现实时运行的语音分离网络,同时在盲源分离和目标语音提取任务上超越现有流式模型。该网络在时频域运作,通过堆叠的全连接层交替处理通道与频率维度,并对每个频率分量独立使用卷积层处理时间序列。实验表明,经过混合精度量化感知训练(QAT)的模型可在GAP9处理器上以6毫秒音频块实现实时处理,相比先前模型运行时间减少3.5至4倍。

原文摘要 · Abstract (English)

Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelerators designed for hearables, due to their limited compute capabilities. We present TF-MLPNet, the first speech separation network capable of running in real-time on such low-power accelerators while outperforming existing streaming models for blind speech separation and target speech extraction. Our network operates in the time-frequency domain, processing frequency sequences with stacks of fully connected layers that alternate along the channel and frequency dimensions, and independently processing the time sequence at each frequency bin using convolutional layers. Results show that our mixed-precision quantization-aware trained (QAT) model can process 6 ms audio chunks in real-time on the GAP9 processor, achieving a 3.5-4x runtime reduction compared to prior speech separation models.

语音分离实时系统低功耗可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。