arXiv:2409.12416eess.AScs.SD2024-09被引 3

用时频联合特征提升语音去削波效果,低信噪比下仍表现稳健。

Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features

  • 融合复数谱图与可学习时域特征,构建时频域联合建模
  • 在VoiceBank-DEMAND和DNS数据集上均超越当前最优模型
  • 保留未削波部分,避免仅用频谱导致的信号退化

我们提出一种基于Transformer的语音去削波模型,可在广泛的输入信噪比(SDR)条件下有效恢复削波信号。尽管近期基于时域深度神经网络(DNN)的去削波方法已优于传统手工设计及基于谱图的DNN方法,但在低SDR输入下仍存在性能瓶颈。为此,我们采用在时频(TF)域操作的Transformer架构。虽然该架构在低SDR语音增强任务中表现出色,但对时域失真如削波并不理想。为克服基于谱图DNN的局限性,我们设计了一个额外的卷积模块,直接从时域波形中提取时序特征。通过联合分析复数谱图与学习到的时序特征,模型在高、低SDR输入下均实现性能提升。所提方法在处理过程中保留未削波语音部分,防止仅使用频谱信息时常见的信号劣化。在VoiceBank-DEMAND与DNS挑战数据集上的评估表明,该模型在各类指标上持续优于当前最先进(SOTA)的去削波模型,展现出鲁棒性与泛化能力。

原文摘要 · Abstract (English)

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have outperformed traditional handcrafted and spectrogram-based DNN approaches, they still struggle with low-SDR inputs. To address this, we incorporate a transformer-based architecture that operates in the time-frequency (TF) domain. The TF-transformer architecture has demonstrated remarkable performance in the speech enhancement task for low-SDR signals but cannot be optimal for the time-domain artifact like clipping. To overcome the limitations of spectrogram-based DNNs, we design an extra convolutional block that directly extracts temporal features from time-domain waveforms. The joint analysis of complex spectrogram and learned temporal features allows the model to improve performance on both high- and low-SDR inputs. Our approach also preserves the unclipped portions of the speech signal during processing, preventing degradation typically seen when only spectral information is used. In evaluations on the VoiceBank-DEMAND and DNS challenge datasets, the proposed model consistently outperformed state-of-the-art (SOTA) declipping models across various metrics, demonstrating its robustness and generalizability.

语音去削波Transformer时频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。