arXiv:2508.01847eess.AScs.LG2025-08被引 2

让语音增强模型在推理时动态适应新噪声环境。

Test-Time Training for Speech Enhancement

  • 用自监督任务引导模型在测试时实时优化。
  • 真实和合成数据集上均显著提升语音质量。
  • 适合需要应对未知噪声的实时语音系统。

本文提出一种将测试时训练(TTT)应用于语音增强的新方法,以应对不可预测的噪声条件和领域偏移问题。该方法采用Y型架构,将主语音增强任务与自监督辅助任务结合。模型在推理阶段通过优化噪声增强信号重建或掩码谱图预测等自监督任务,动态适应新领域,无需标注数据。文中还设计了多种TTT策略,在适应性与效率间提供权衡。在合成和真实世界数据集上的评估显示,各项语音质量指标均持续优于基线模型。本工作验证了TTT在语音增强中的有效性,为未来自适应、鲁棒语音处理研究提供了重要启示。

原文摘要 · Abstract (English)

This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhancement task with a self-supervised auxiliary task in a Y-shaped architecture. The model dynamically adapts to new domains during inference time by optimizing the proposed self-supervised tasks like noise-augmented signal reconstruction or masked spectrogram prediction, bypassing the need for labeled data. We further introduce various TTT strategies offering a trade-off between adaptation and efficiency. Evaluations across synthetic and real-world datasets show consistent improvements across speech quality metrics, outperforming the baseline model. This work highlights the effectiveness of TTT in speech enhancement, providing insights for future research in adaptive and robust speech processing.

语音增强测试时训练自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。