arXiv:2601.14770eess.AS2026-01中稿 · ICASSP 2026被引 2

通过增强掩码二值性,让语音增强模型在未知环境更稳定。

Test-Time Adaptation For Speech Enhancement Via Mask Polarization

  • 用瓦斯赫斯坦距离比较分布,恢复掩码的二峰特性
  • 在多种场景下提升性能,且无需额外参数
  • 适合边缘设备部署,对资源敏感的语音应用很实用

将语音增强(SE)模型适应到未见环境对实际部署至关重要,但测试时自适应(TTA)在SE领域仍研究不足,主要因缺乏对模型在域偏移下退化机制的理解。我们观察到,基于掩码的SE模型在域偏移下会丧失置信度,预测掩码趋于平滑,失去对语音的明确保留与噪声的有效抑制。基于此洞察,我们提出掩码极化(MPol),一种轻量级TTA方法,通过瓦斯赫斯坦距离进行分布对比,恢复掩码的双峰特性。MPol仅需原训练模型参数,无额外参数,适合资源受限的边缘部署。在多种域偏移和模型架构下的实验表明,MPol实现了稳定且与复杂方法相当的性能提升。

原文摘要 · Abstract (English)

Adapting speech enhancement (SE) models to unseen environments is crucial for practical deployments, yet test-time adaptation (TTA) for SE remains largely under-explored due to a lack of understanding of how SE models degrade under domain shifts. We observe that mask-based SE models lose confidence under domain shifts, with predicted masks becoming flattened and losing decisive speech preservation and noise suppression. Based on this insight, we propose mask polarization (MPol), a lightweight TTA method that restores mask bimodality through distribution comparison using the Wasserstein distance. MPol requires no additional parameters beyond the trained model, making it suitable for resource-constrained edge deployments. Experimental results across diverse domain shifts and architectures demonstrate that MPol achieves very consistent gains that are competitive with significantly more complex approaches.

语音增强测试时适应掩码极化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。