2025 URGENT语音增强挑战赛探索多语言、多畸变场景下的通用语音增强方法。
Interspeech 2025 URGENT Speech Enhancement Challenge
- 采用混合与判别式模型,评估不同方法在多种语音畸变下的表现。
- 主观评测显示部分生成式/混合模型优于顶尖判别模型。
- 发现纯生成模型存在语言依赖性,对跨语言泛化带来挑战。
近年来,针对多种语音畸变和录音条件的通用语音增强(SE)研究日益增多。URGENT挑战系列旨在通过涵盖广泛畸变类型、增加数据多样性并引入丰富评估指标,推动通用语音增强的发展。本文介绍了第二期的Interspeech 2025 URGENT挑战赛,重点探索此前关注较少的方面:语言依赖性、更广泛的畸变类型泛化能力、数据可扩展性,以及使用带噪训练数据的有效性。共收到32份提交,最佳系统采用判别式模型,多数竞争性方案为混合方法。分析揭示两个关键发现:(i) 主观评测中,某些生成式或混合方法优于顶级判别模型;(ii) 纯生成式语音增强模型表现出明显的语言依赖性。
原文摘要 · Abstract (English)
There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embracing a broad range of distortion types, increasing data diversity, and incorporating extensive evaluation metrics. This work introduces the Interspeech 2025 URGENT Challenge, the second edition of the series, to explore several aspects that have received limited attention so far: language dependency, universality for more distortion types, data scalability, and the effectiveness of using noisy training data. We received 32 submissions, where the best system uses a discriminative model, while most other competitive ones are hybrid methods. Analysis reveals some key findings: (i) some generative or hybrid approaches are preferred in subjective evaluations over the top discriminative model, and (ii) purely generative SE models can exhibit language dependency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。