arXiv:2505.13843eess.AScs.SD2025-05中稿 · interspeech 2025被引 2

用分层建模语义与声学特征,提升复杂环境下的语音增强效果。

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

  • 分步分解语义与声学信息,通过因子化编解码器与扩散模型实现
  • 在复杂噪声下语音质量优于当前最优方法,且提升下游TTS表现
  • 适合需要高质量语音输出的语音合成、降噪应用

现有语音增强方法多通过直接估计时频掩码或频谱来恢复干净语音,但常忽略语音信号中固有的语义内容与声学细节,导致下游任务性能受限,尤其在复杂声学环境中表现下降。为此,本文提出一种基于语义信息的分层语音增强方法,结合因子化编解码器与扩散模型,分步建模语音的语义与声学属性。该方法在挑战性声学场景下仍能实现更鲁棒的语音恢复,并显著提升下游语音合成(TTS)任务的表现。实验表明,该算法在语音质量上超越当前最优基线,同时有效改善噪声环境下TTS的生成效果。

原文摘要 · Abstract (English)

Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.

语音增强扩散模型语义建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。