arXiv:2506.01611eess.AScs.SD2025-06中稿 · Interspeech 2025被引 14

剖析语音增强系统开发中的数据清洗与评估盲区。

Lessons Learned from the URGENT 2024 Speech Enhancement Challenge

  • 揭示语音数据带宽标注与实际不符、高质量语料仍存标签噪声。
  • 发现现有系统难以应对强混响、重叠语音等极端场景,且缺乏难度评估标准。
  • 主张融合多维度指标以更贴近人类听感,提升评估可靠性。

URGENT 2024 语音增强挑战赛旨在推动具备高度通用性、鲁棒性和泛化能力的语音增强技术,涵盖更广泛的任务定义、大规模跨域数据集及全面的评估指标。基于该挑战赛结果,本文深入分析了语音增强系统开发中两个被忽视的关键问题:数据清洗与评估指标。我们指出传统语音增强流程中存在的若干被忽略问题:(1)声明音频带宽与实际有效带宽不一致,即便在多种“高质量”语音语料库中也存在标签噪声;(2)现有系统难以应对最严峻条件(如语音重叠、强噪声/混响),且缺乏对语音样本难度的有效衡量手段;(3)必须结合多维度指标进行综合评估,才能与人类听觉判断具有良好相关性。本文期望能启发未来更优语音增强流水线的设计。

原文摘要 · Abstract (English)

The URGENT 2024 Challenge aims to foster speech enhancement (SE) techniques with great universality, robustness, and generalizability, featuring a broader task definition, large-scale multi-domain data, and comprehensive evaluation metrics. Nourished by the challenge outcomes, this paper presents an in-depth analysis of two key, yet understudied, issues in SE system development: data cleaning and evaluation metrics. We highlight several overlooked problems in traditional SE pipelines: (1) mismatches between declared and effective audio bandwidths, along with label noise even in various "high-quality" speech corpora; (2) lack of both effective SE systems to conquer the hardest conditions (e.g., speech overlap, strong noise / reverberation) and reliable measure of speech sample difficulty; (3) importance of combining multifaceted metrics for a comprehensive evaluation correlating well with human judgment. We hope that this endeavor can inspire improved SE pipeline designs in the future.

语音增强数据清洗评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。