arXiv:2503.04713eess.AScs.AI2025-03EMNLP被引 25

构建首个大规模语音风格标签数据集,提升文本转语音的风格一致性与自然度。

Scaling Rich Style-Prompted Text-to-Speech Datasets

  • 用文本、语音嵌入和音频语言模型自动标注59种丰富风格标签
  • 在342小时人工标注+2427小时自动生成数据上训练,风格一致性提升7.9%
  • 适合语音合成、风格控制研究者使用,开源可复现

我们提出Paralinguistic Speech Captions(ParaSpeechCaps),一个大规模语音风格标注数据集,对语音语句进行丰富风格描述。现有大规模数据集仅包含基础标签(如低音、慢速、大声),而小规模人工标注数据虽有抽象风格标签(如喉音、鼻音、痛苦),但难以扩展。本文首次结合现成的文本与语音嵌入模型、分类器及音频语言模型,实现丰富风格标签的自动化扩展。ParaSpeechCaps包含59种风格标签,涵盖说话人级内在特征与话语级情境特征,共含342小时人工标注数据(PSC-Base)和2427小时自动生成数据(PSC-Scaled)。我们在开源风格提示文本转语音模型Parler-TTS上微调该数据集,相比现有最佳基线(融合多个已有风格数据集),风格一致性提升7.9%(一致性MOS),语音自然度提升15.5%(自然度MOS)。通过消融实验验证了数据设计选择的有效性,为后续研究奠定基础。数据集、模型与代码已开源。

原文摘要 · Abstract (English)

We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g. low-pitched, slow, loud). We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time. ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags. It consists of 342 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled). We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets. We ablate several of our dataset design choices to lay the foundation for future work in this space. Our dataset, models and code are released at https://github.com/ajd12342/paraspeechcaps .

语音合成风格控制数据集TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。