arXiv:2509.04280eess.AS2025-09被引 2

无需标签数据,实时适配语音增强模型应对真实环境变化。

Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation

  • 用预训练语音表征进行隐空间去噪,线性变换逼近干净语音特征。
  • 在多语言、多说话人等域偏移下,性能优于现有方法。
  • 适合部署在噪声复杂、语音多样的真实场景中使用。

基于深度学习的语音增强模型在测试分布与训练条件一致时表现优异,但在实际环境中因领域偏移导致性能下降。为此,我们提出LaDen(隐空间去噪),首个专为语音增强设计的测试时自适应方法。该方法利用强大的预训练语音表征,通过噪声嵌入的线性变换实现隐空间去噪,近似还原干净语音表示。此变换具备跨领域泛化能力,可在无标签目标域上生成有效伪标签,进而实现模型在多样化声学环境中的测试时自适应。我们构建了涵盖多个数据集及多种领域偏移(如噪声类型、说话人特征、语言变化)的综合基准。大量实验表明,LaDen在感知指标上持续优于基线方法,尤其在说话人和语言域偏移下表现突出。

原文摘要 · Abstract (English)

Deep learning-based speech enhancement models achieve remarkable performance when test distributions match training conditions, but often degrade when deployed in unpredictable real-world environments with domain shifts. To address this challenge, we present LaDen (latent denoising), the first test-time adaptation method specifically designed for speech enhancement. Our approach leverages powerful pre-trained speech representations to perform latent denoising, approximating clean speech representations through a linear transformation of noisy embeddings. We show that this transformation generalizes well across domains, enabling effective pseudo-labeling for target domains without labeled target data. The resulting pseudo-labels enable effective test-time adaptation of speech enhancement models across diverse acoustic environments. We propose a comprehensive benchmark spanning multiple datasets with various domain shifts, including changes in noise types, speaker characteristics, and languages. Our extensive experiments demonstrate that LaDen consistently outperforms baseline methods across perceptual metrics, particularly for speaker and language domain shifts.

语音增强测试时自适应域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。