用噪声不变语义蒸馏,让语音增强更准确不乱编。
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

- 从噪声语音中学习鲁棒条件编码,同时蒸馏声学与语义目标。
- 低信噪比和混响环境下,语言一致性提升显著,音质仍优秀。
- 适合需要高保真语音生成的场景,如降噪通话、语音助手。
基于语言模型(LM)的语音增强能生成自然语音,但在强噪声下常因条件不可靠导致输出语义错误。本文提出L3-SE,一种噪声不变的声学-语义蒸馏框架,通过联合蒸馏两个互补的清晰语音目标——声学目标用于重建保真度,语义目标用于语言一致性,从含噪语音中学习噪声不变的条件编码器。利用获得的噪声不变声学-语义表征,驱动解码器仅的自回归语言模型预测干净声学标记,并通过高保真编码器还原为增强语音。该编码器基于可学习加权WavLM层表示构建。实验表明,所提方法在语言一致性指标上持续优于现有基线,尤其在低信噪比(low-SNR)和混响条件下提升明显,同时保持竞争力的感知质量。音频样本见 https://max1wz.github.io/L3-SE-Demo-Page/。完整代码将在论文录用后公开。
原文摘要 · Abstract (English)
Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptually plausible yet linguistically incorrect outputs. To address this issue, we propose L3-SE, a noise-invariant acoustic-semantic distillation framework for reducing linguistic hallucination in LM-based SE. The proposed method learns a noise-invariant conditioning encoder from noisy speech by jointly distilling two complementary clean-speech targets: an acoustic target for reconstruction fidelity and a semantic target for linguistic consistency. The resulting noise-invariant acoustic-semantic representations are used to condition a decoder-only autoregressive language model, which predicts clean acoustic tokens that are decoded into enhanced speech. To support high-quality generation, we further employ a high-fidelity codec built on learnable weighted WavLM layer representations as the discrete acoustic interface. By improving the reliability of conditioning under adverse conditions, the proposed framework substantially reduces hallucination and improves content faithfulness. Experiments show that the proposed method consistently outperforms prior LM-based speech enhancement baselines on linguistic consistency metrics, with especially clear gains under low-SNR and reverberant conditions, while maintaining competitive perceptual quality. Audio samples are available at https://max1wz.github.io/L3-SE-Demo-Page/. The complete source code will be released after the manuscript is accepted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。