arXiv:2603.14328cs.SDeess.AS2026-03被引 3

构建跨口音语音合成评估数据集,揭示音色与口音的感知关联。

CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents

  • 采集24套系统生成的4000个重合成与TTS样本,覆盖10种口音
  • 25名听者对自然度、音色相似性、口音相似性评分,获19600条标注
  • 发现同口音听众存在感知偏见,客观指标可预测主观评价

我们提出CodecMOS-Accent数据集,一个用于评估神经音频编解码器(NAC)模型及基于其训练的大语言模型(LLM)文本到语音(TTS)模型的均值意见分数(MOS)基准,尤其针对非标准语音如口音语音。该数据集包含来自24个系统的4,000个编解码重合成和TTS样本,涵盖32位讲者,覆盖十种口音。通过大规模主观测试,从25名听者处收集了19,600条标注,评估维度包括自然度、说话人相似性和口音相似性。该数据集不仅反映了近期语音合成系统性能的最新状况,还揭示了说话人与口音相似性之间的紧密关系、客观指标的预测能力,以及当听者与说话人同口音时存在的感知偏见。该数据集有望推动更以人为本的NAC和口音化TTS研究。

原文摘要 · Abstract (English)

We present the CodecMOS-Accent dataset, a mean opinion score (MOS) benchmark designed to evaluate neural audio codec (NAC) models and the large language model (LLM)-based text-to-speech (TTS) models trained upon them, especially across non-standard speech like accented speech. The dataset comprises 4,000 codec resynthesis and TTS samples from 24 systems, featuring 32 speakers spanning ten accents. A large-scale subjective test was conducted to collect 19,600 annotations from 25 listeners across three dimensions: naturalness, speaker similarity, and accent similarity. This dataset does not only represent an up-to-date study of recent speech synthesis system performance but reveals insights including a tight relationship between speaker and accent similarity, the predictive power of objective metrics, and a perceptual bias when listeners share the same accent with the speaker. This dataset is expected to foster research on more human-centric evaluation for NAC and accented TTS.

语音合成口音评估主观评测神经编解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。