arXiv:2607.08645cs.SD2026-07

量化压缩让双耳语音增强模型更轻量,仍保持高音质。

It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement

论文配图:It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement
图 1 · 摘自论文原文
  • 对双耳语音增强模型进行量化压缩,降低计算与内存开销。
  • 即使中间掩码估计被量化,空间滤波阶段仍能补偿误差。
  • 最终模型仅需4.65 MMAC/s和0.177 MB,适合移动端部署。

基于神经网络的多通道语音增强系统虽性能出色,但计算与内存需求限制了其在资源受限设备上的部署。本文研究TANGO——一种结合神经掩码估计与空间滤波的分布式双耳语音增强系统的低精度推理。评估了后训练量化与量化感知训练对神经组件的影响,并分析量化误差在掩码估计器中如何传播至下游空间滤波阶段。结果表明,尽管量化会降低中间掩码估计精度,但空间滤波阶段可有效补偿大部分误差。利用此鲁棒性,我们将TANGO简化为MN-TANGO,显著降低模型尺寸与计算复杂度,同时保持相近的最终性能。通过结合INT8权重与激活量化、ERB压缩及分组循环层,最紧凑的MN-TANGO实现4.65 MMAC/s和0.177 MB。

原文摘要 · Abstract (English)

Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-constrained devices. This paper investigates low-precision inference for TANGO, a hybrid distributed binaural speech enhancement system combining neural mask estimation with spatial filtering. We evaluate post-training quantization and quantization-aware training for the neural components, and analyze how quantization errors in the mask estimators propagate through the downstream spatial filtering stage. Our analysis shows that, although quantization degrades intermediate mask estimates, the spatial filtering stage compensates for most quantization-induced errors. Leveraging this robustness, we simplify TANGO into MN-TANGO, reducing both model size and computational complexity while maintaining comparable final performance. By combining INT8 weight-and-activation quantization with ERB compression and grouped recurrent layers, the most compact MN-TANGO reaches 4.65 MMAC/s and 0.177 MB.

语音增强模型量化双耳处理轻量部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。