arXiv:2505.15061cs.SDeess.AS2025-05被引 16

SHEET工具箱可快速评估语音质量,提升研究效率。

SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit

  • 基于深度神经网络预测人类评分的语音质量评估工具
  • 在BVCC和NISQA数据集上验证,性能优于原SSL-MOS
  • 支持多数据集与预训练模型,适合语音质量研究者使用

我们提出SHEET,一个多功能开源工具箱,旨在加速主观语音质量评估(SSQA)研究。SHEET即语音人类评估估计工具箱,聚焦于基于数据驱动的深度神经网络模型,用于预测语音样本的人类标注质量分数。该工具箱提供完整的训练与评估脚本,支持多数据集和多模型,并可通过Torch Hub和HuggingFace Spaces获取预训练模型。为展示其能力,我们在一系列语音自监督学习(SSL)模型上重新评估了广泛使用的SSL-MOS模型。实验在两个代表性SSQA数据集BVCC和NISQA上进行,结果识别出最优语音SSL模型,其性能超越原始SSL-MOS实现,并达到当前最先进方法水平。

原文摘要 · Abstract (English)

We introduce SHEET, a multi-purpose open-source toolkit designed to accelerate subjective speech quality assessment (SSQA) research. SHEET stands for the Speech Human Evaluation Estimation Toolkit, which focuses on data-driven deep neural network-based models trained to predict human-labeled quality scores of speech samples. SHEET provides comprehensive training and evaluation scripts, multi-dataset and multi-model support, as well as pre-trained models accessible via Torch Hub and HuggingFace Spaces. To demonstrate its capabilities, we re-evaluated SSL-MOS, a speech self-supervised learning (SSL)-based SSQA model widely used in recent scientific papers, on an extensive list of speech SSL models. Experiments were conducted on two representative SSQA datasets named BVCC and NISQA, and we identified the optimal speech SSL model, whose performance surpassed the original SSL-MOS implementation and was comparable to state-of-the-art methods.

语音评估开源工具深度学习自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。