arXiv:2601.14516eess.AScs.SD2026-01中稿 · presentation at IC…被引 1

通过共享语音表征联合训练降噪与语音还原,提升嘈杂环境下的还原效果。

Towards noise-robust speech inversion through multi-task learning with speech enhancement

  • 共享自监督语音表征,联合优化降噪与语音还原模块。
  • 在-5dB信噪比下,不同噪声场景下相关性提升超38%以上。
  • 适合需要抗噪语音还原的实时应用,如语音通信与听障辅助。

近期研究证明自监督学习(SSL)语音表征在语音还原(SI)任务中的有效性。然而,在真实场景中,背景噪声普遍存在,仍使SI应用面临挑战。本文提出一种统一框架,通过共享的基于SSL的语音表征,将语音增强(SE)与语音还原(SI)模型集成。在此框架中,SSL模型不仅用于支持SE模块抑制噪声,还生成对SI任务更具信息量的表示,使两个模块在联合训练中相互受益。在-5 dB信噪比条件下,该方法在人声干扰噪声下相比基线相对提升80.95%,在非人声干扰噪声下提升38.98%,以所有估计参数间的平均皮尔逊相关系数衡量。

原文摘要 · Abstract (English)

Recent studies demonstrate the effectiveness of Self Supervised Learning (SSL) speech representations for Speech Inversion (SI). However, applying SI in real-world scenarios remains challenging due to the pervasive presence of background noise. We propose a unified framework that integrates Speech Enhancement (SE) and SI models through shared SSL-based speech representations. In this framework, the SSL model is trained not only to support the SE module in suppressing noise but also to produce representations that are more informative for the SI task, allowing both modules to benefit from joint training. At a Signal-to-Noise Ratio of -5 db, our method for the SI task achieves relative improvements over the baseline of 80.95% under babble noise and 38.98% under non-babble noise, as measured by the average Pearson product-moment correlation across all estimated parameters.

语音还原降噪自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。