arXiv:2507.01805cs.SDeess.AS2025-07中稿 · Interspeech 2025被引 1

首个西班牙语语音自然度自动评估数据集,助力提升语音合成质量预测精度。

A Dataset for Automatic Assessment of TTS Quality in Spanish

  • 构建4326个样本的西班牙语语音数据集,涵盖52种语音系统与真人发音。
  • 基于主观测试训练的模型在5分制自然度评分上误差仅0.8分。
  • 适用于西班牙语语音合成研究,尤其适合优化自然度评估系统。

本文致力于构建首个用于西班牙语文本转语音(TTS)系统自动评估的数据集,旨在提升自然度预测模型的准确性。该数据集包含来自52种不同TTS系统和真人发音的4,326个音频样本,是目前西班牙语领域首个同类数据集。音频标签基于ITU-T Rec. P.807标准设计的主观测试,由92名参与者完成。通过训练自动自然度预测模型验证数据集有效性,采用两种方法:微调原始针对英语训练的模型,以及在冻结的自监督语音模型基础上训练小型下游网络。模型在五点平均意见分(MOS)尺度上的平均绝对误差为0.8。进一步分析表明,该数据集具备高质量与多样性,具有推动西班牙语语音合成研究的潜力。

原文摘要 · Abstract (English)

This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52 different TTS systems and human voices and is, up to our knowledge, the first of its kind in Spanish. To label the audios, a subjective test was designed based on the ITU-T Rec. P.807 standard and completed by 92 participants. Furthermore, the utility of the collected dataset was validated by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model originally trained for English, and training small downstream networks on top of frozen self-supervised speech models. Our models achieve a mean absolute error of 0.8 on a five-point MOS scale. Further analysis demonstrates the quality and diversity of the developed dataset, and its potential to advance TTS research in Spanish.

语音合成自然度评估西班牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。