arXiv:2502.06839eess.AScs.AI2025-02被引 2

用少量声学信息训练语音去混响模型,效果更稳定。

A Hybrid Model for Weakly-Supervised Speech Dereverberation

  • 仅需混响时间等有限声学信息训练,不依赖成对干/湿语音数据。
  • 在多个客观指标上表现优于当前最优方法,性能更一致。
  • 适合缺乏标注数据、追求鲁棒性的语音增强应用场景。

本文提出一种新型训练策略,用于弱监督语音去混响。现有算法多依赖成对的干/湿语音数据(难以获取),或使用目标度量指标,但这些指标可能无法充分捕捉混响特性,导致非目标指标表现差。本文方法仅利用有限的声学信息(如混响时间RT60)训练去混响系统。通过生成房间冲激响应对输出进行重合成,并与原始混响语音对比,构建新型混响匹配损失,替代传统目标度量。推理时仅使用训练好的去混响模型。实验表明,该方法在多种语音去混响常用客观指标上均取得更一致且优越的性能,优于当前最先进方法。

原文摘要 · Abstract (English)

This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is difficult to obtain, or on target metrics that may not adequately capture reverberation characteristics and can lead to poor results on non-target metrics. Our approach uses limited acoustic information, like the reverberation time (RT60), to train a dereverberation system. The system's output is resynthesized using a generated room impulse response and compared with the original reverberant speech, providing a novel reverberation matching loss replacing the standard target metrics. During inference, only the trained dereverberation model is used. Experimental results demonstrate that our method achieves more consistent performance across various objective metrics used in speech dereverberation than the state-of-the-art.

语音去混响弱监督声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。