arXiv:2505.21156cs.SDcs.AI2025-05中稿 · Interspeech 2025

用模型自身编码器做损失函数,提升语音增强效果

Model as Loss: A Self-Consistent Training Paradigm

  • 用同一模型的编码器作为损失函数,实现自洽训练
  • 在标准数据集上优于预训练特征损失,提升听感质量
  • 适合追求高质量语音增强与泛化能力的研究者

传统语音增强方法依赖人工设计的损失函数(如时域或频域损失)或深度特征损失(如WavLM或wav2vec),常无法捕捉对性能至关重要的细微信号特性。为此,我们提出Model as Loss新训练范式,利用同一模型的编码器作为损失函数来指导训练。该范式借助编码器的任务特定特征空间,使解码器输出与纯净语音的感知和任务相关特征保持一致。通过使用编码器学习到的特征作为损失函数,该框架强化了纯净参考语音与增强模型输出之间的自洽性。实验表明,该方法在标准语音增强基准测试中优于预训练深度特征损失,在听感质量和跨域泛化能力上均有提升。

原文摘要 · Abstract (English)

Conventional methods for speech enhancement rely on handcrafted loss functions (e.g., time or frequency domain losses) or deep feature losses (e.g., using WavLM or wav2vec), which often fail to capture subtle signal properties essential for optimal performance. To address this, we propose Model as Loss, a novel training paradigm that utilizes the encoder from the same model as a loss function to guide the training. The Model as Loss paradigm leverages the encoder's task-specific feature space, optimizing the decoder to produce output consistent with perceptual and task-relevant characteristics of the clean signal. By using the encoder's learned features as a loss function, this framework enforces self-consistency between the clean reference speech and the enhanced model output. Our approach outperforms pre-trained deep feature losses on standard speech enhancement benchmarks, offering better perceptual quality and robust generalization to both in-domain and out-of-domain datasets.

语音增强自洽训练深度特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。