arXiv:2511.13300eess.AScs.AI2025-11AAAI被引 7

用预训练模型的发音规律抑制语音增强中的幻觉问题。

PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement

  • 利用WavLM的发音先验作为约束,防止生成错误内容。
  • 在严重噪声下,语言和声学幻觉显著降低。
  • 适合追求高保真语音重建的研究者与工程师。

生成式模型在语音增强(SE)中表现出色,感知质量优于传统判别式方法。然而,现有生成式SE方法在强噪声下常出现幻觉,导致语音内容错误或说话人特征不一致,分别称为语言和声学幻觉。我们指出,语言幻觉源于模型未能约束有效的发音结构,是更根本的挑战。尽管语言模型(LMs)可通过离散标记分布捕捉语音结构,但现有方法难以从噪声污染的表示中学习,导致先验信息被污染并引发幻觉。为此,我们提出音系锚定语音增强器(PASE),一种利用预训练WavLM模型内在发音先验的生成式SE框架。首先,通过表示蒸馏将WavLM改造为去噪专家,以清洁其最终层特征;基于模型内在的发音先验,该过程实现鲁棒去噪并最小化语言幻觉。为进一步减少声学幻觉,我们采用双流表示训练声码器:高层语音表示提供清晰的语言内容,低层声学表示保留说话人身份与语调。实验表明,PASE不仅在感知质量上超越现有判别式模型,且显著优于先前生成式模型,同时大幅降低语言与声学幻觉。

原文摘要 · Abstract (English)

Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations.

语音增强生成模型幻觉抑制发音先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。