arXiv:2509.21185eess.AS2025-09

混合实数与复数神经网络提升语音增强效果并降低计算量

Hybrid Real- and Complex-Valued Neural Network Architecture for Speech Enhancement

  • 用实数掩码分支+复数修正分支,通过域转换耦合
  • 在四种信噪比下均提升语音清晰度与质量
  • 显著减少运算量,适合小模型部署

本文研究了用于单声道语音增强的混合实数与复数神经网络。虽然复数模型能原生处理时频表示,但通常增加计算成本,且在小模型场景下效率较低。为此,我们提出一种参数匹配的混合架构:将实数幅度掩码分支与复数加性修正分支通过瓶颈处的域转换函数耦合。该方法应用于卷积去噪自编码器和卷积-循环网络架构。在四个信噪比下的平均性能显示,混合模型相比基线显著提升语音可懂度与质量,同时大幅降低运算量。

原文摘要 · Abstract (English)

This paper investigates hybrid real- and complex-valued neural networks for monaural speech enhancement. While complex-valued models can process time-frequency representations natively, they often increase computational cost and can be inefficient in small-model regimes. We therefore study a matched-parameter hybrid architecture that combines a real-valued magnitude-mask branch with a complex-valued additive correction branch, coupled via domain conversion functions at the bottleneck. The approach is applied to convolutional denoising autoencoder and convolutional-recurrent network architectures. Averaged over four SNRs, the hybrid models improve intelligibility and quality over the considered counterparts, while substantially reducing the number of operations.

语音增强混合网络复数神经网络低复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。