arXiv:2506.02863eess.AScs.AI2025-06被引 22

构建首个覆盖多风格的语音合成基准,支持真实场景应用。

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

  • 设计多任务语音合成数据集,包含百万级标注音频
  • 实现高保真多样风格语音生成,支持声音事件与情感控制
  • 适合语音系统开发、数字人与智能客服研究者

近期生成式人工智能显著推动了风格化文本到语音合成(CapTTS)的发展。然而,由于缺乏标准化、全面的数据集以及对下游任务的研究不足,将CapTTS应用于实际场景仍面临挑战。为此,我们提出CapSpeech,一个面向多种CapTTS相关任务的新基准,包括带音效的风格化语音合成(CapTTS-SE)、带口音的语音合成(AccCapTTS)、带情绪的语音合成(EmoCapTTS)以及聊天机器人语音合成(AgentTTS)。该数据集包含超过1000万条机器标注的音视频对和近36万条人工标注的音视频对。此外,我们还专门收集并录制了由专业配音演员和音频工程师完成的两个新数据集,分别用于AgentTTS和CapTTS-SE任务。基于该数据集,我们使用自回归与非自回归模型进行了全面实验,结果表明在多种语音风格下均能实现高保真、高可懂度的语音合成。据我们所知,CapSpeech是目前规模最大的提供完整标注的CapTTS相关任务数据集,其实验结果为开发高效CapTTS系统提供了重要参考。

原文摘要 · Abstract (English)

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.

语音合成风格控制多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。