arXiv:2506.23874eess.AScs.SD2025-06被引 3

用配对比较提升语音增强系统排名,省去大量人工评分

URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition

  • 通过对比成对语音输出,预测系统质量相对排名
  • 在有限数据下仍优于现有最佳方法,跨多个测试集表现稳定
  • 适合语音增强竞赛评估,尤其资源受限场景

平均意见分(MOS)是语音质量评估的核心,但获取需大量人工标注。尽管已有如DNSMOS和UTMOS等深度学习模型可预测MOS以避免此问题,但常因训练数据不足而表现不佳。鉴于语音增强系统比较更关注可靠排序而非绝对分数,我们提出URGENT-PK,一种基于配对比较的新型排名方法。该模型以同源增强语音对为输入,预测其相对质量排序。这种配对范式能高效利用有限训练数据,因多系统间所有配对组合可构成一个训练实例。在多个公开测试集上的实验表明,尽管网络结构简单且训练数据有限,URGENT-PK在系统级排序性能上仍优于当前最先进基线。

原文摘要 · Abstract (English)

The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK's superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.

语音增强排名模型配对比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。