arXiv:2510.01130eess.AS2025-10中稿 · IEEE TASLP被引 1

提出可学习图频谱表示,提升单声道语音增强效果

Learning Time-Graph Frequency Representation for Monaural Speech Enhancement

  • 用1D卷积构建可学习图拓扑,替代固定结构
  • 避免矩阵求逆,解决数值误差与不稳定性问题
  • 适合需要高稳定性语音增强的场景

图傅里叶变换(GFT)在语音增强中表现出良好性能。然而,现有基于GFT的方法通常采用固定图拓扑构建图傅里叶基,缺乏自适应性与灵活性。此外,基于奇异值分解(GFT-SVD)和特征向量分解(GFT-EVD)的GFT方法会因矩阵求逆引入数值误差与不稳定性。为此,本文提出一种简单而有效的可学习GFT-SVD框架用于语音增强。具体地,利用图移位算子构建可学习图拓扑,并通过1维卷积神经层生成可学习的图傅里叶基,基于奇异值矩阵定义。该方法无需矩阵求逆,从而避免相关数值误差与稳定性问题。

原文摘要 · Abstract (English)

The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed graph topologies to build the graph Fourier basis, whose the representation lacks the adaptively and flexibility. In addition, they suffer from the numerical errors and instability introduced by matrix inversion in GFT based on both Singular Value Decomposition (GFT-SVD) and Eigen Vector Decomposition (GFT-EVD). Motivated by these limitations, this paper propose a simple yet effective learnable GFT-SVD framework for speech enhancement. Specifically, we leverage graph shift operators to construct a learnable graph topology and define a learnable graph Fourier basis by the singular value matrices using 1-D convolution (Conv-1D) neural layer. This eliminates the need for matrix inversion, thereby avoiding the associated numerical errors and stability problem.

语音增强图神经网络可学习表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。