arXiv:2411.03715cs.SDeess.AS2024-11中稿 · Transactions on Au…被引 42

构建SSQA评估数据集,发现数据多样性比数量更重要

MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models

  • 构建包含8个训练集和17个测试集的MOS-Bench数据集
  • 多数据集合并比单一数据源更能提升模型泛化能力
  • 强调数据多样性对跨域评估的关键作用,适合语音质量研究者

本文研究主观语音质量评估(SSQA)任务,即预测语音的感知质量。随着深度神经网络的发展,SSQA已取得显著进展,并广泛用于评估语音生成系统。然而,当前SSQA模型在域外(OOD)场景下的泛化能力不足,尚未被充分研究。为此,本文提出MOS-Bench,一个包含8个训练集和17个测试集的多样化SSQA数据集集合。通过大量实验,我们揭示了现有模型在跨域场景下的泛化挑战,并评估了多数据集训练的有效性,对比了简单数据拼接与AlignNet等领域感知方法。结果表明,合并多个训练集是一种简单而有效的解决方案,数据多样性是超越训练规模影响实现稳健泛化的关键因素。

原文摘要 · Abstract (English)

In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neural network models, SSQA has greatly advanced and has been widely applied in scientific papers to evaluate speech generation systems. Nonetheless, the insufficient out-of-domain (OOD) generalization ability of current SSQA models is underexplored and often overlooked by researchers. To study this problem systematically, we present MOS-Bench, a diverse SSQA dataset collection that currently contains 8 training sets and 17 test sets. Through extensive experiments, we first highlight the OOD generalization challenges of existing models. We then evaluate the efficacy of multiple-dataset training, comparing straightforward data pooling against AlignNet, an existing domain-aware method. We demonstrate that pooling multiple training sets provides a simple yet effective solution, and variation in the data is a key factor for robust generalization beyond training data size.

语音评估泛化能力数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。