用自回归先验在测试时优化语音增强模型,提升噪声不匹配场景下的表现。
Test-time adaptation for speech enhancement with an autoregressive speech prior

- 基于自回归先验约束增强语音分布,实现单句测试时适应。
- 在多个噪声数据集上显著提升语音质量,尤其在训练测试噪声不匹配时。
- 无需标签数据,适合部署在未知噪声环境的实时语音增强系统。
测试时适应(TTA)为在声学条件不匹配的情况下提升语音增强模型性能提供了有前景的方向,且无需目标端标注数据。本文提出一种单句话测试时适应方法,通过使用在神经音频编解码器提取的干净语音潜在表示上训练的自回归先验,对预训练语音增强模型进行正则化。适应过程通过最小化增强语音分布与干净语音先验之间的Kullback-Leibler散度实现。在多个含噪语音数据集上的实验表明,该方法在语音质量上均取得稳定提升,尤其在训练-测试噪声不匹配条件下表现突出。代码与音频示例已公开。
原文摘要 · Abstract (English)
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。