arXiv:2509.04072eess.AScs.CL2025-09ACL被引 1

用小说朗读片段构建新数据集,提升语音合成的情感表现力。

Computational Narrative Understanding for Expressive Text-to-Speech

  • 从小说对白中提取5.3小时富有表现力的语音,标注语调关键词
  • 在新数据集上微调模型,显著提升语音自然度与可懂度
  • 开源数据与代码,适合语音合成与叙事理解研究者使用

近年来文本转语音(TTS)的发展依赖大规模跨领域语音语料库,但人类朗读的有声书,尤其是虚构作品中的表达潜力仍未被充分挖掘。我们发现,这类作品中中性叙述与角色对话的自然交替,蕴含丰富多样的语调线索。基于此,我们构建了大型数据集LibriQuote,包含5.3小时来自角色对白的富有表现力语音,并为每段语音添加上下文伪标签,用于标注说话动词和副词(如“轻声说”),以刻画直接引语的预期表达方式。实验表明,在LibriQuote上微调流匹配模型能显著提升语音的表现力与可懂度;而从头训练自回归TTS模型则增强了其表达能力。在LibriQuote-test上的基准测试显示各系统生成表达性语音的能力差异显著。我们已公开发布数据集、代码与评估资源,以促进复现。音频样本见 https://libriquote.github.io/。

原文摘要 · Abstract (English)

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works, contain rich and diverse prosodic cues arising from the natural alternation between neutral narration and expressive character dialogue. Building from this observation, we introduce LibriQuote, a large-scale 5.3K hours of expressive speech drawn from character quotations. Each quote is supplemented with contextual pseudo-labels for speech verbs and adverbs that characterize the intended delivery of direct speech (e.g., "he whispered softly"). We found that fine-tuning a flow-matching model on LibriQuote yields substantial improvements in expressivity and intelligibility, while training from scratch enhances expressiveness of an autoregressive TTS model. Benchmarking on LibriQuote-test highlights significant variability across systems in generating expressive speech. We publicly release the dataset, code, and evaluation resources to facilitate reproducibility. Audio samples can be found at https://libriquote.github.io/.

语音合成情感表达数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。