arXiv:2509.14946eess.AScs.CL2025-09被引 7

自动生成118小时语音副语言数据,提升语音合成真实感

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

  • 基于自然对话语音自动提取并标注副语言信号
  • 构建含6类118.75小时精确时间戳的SynParaSpeech数据集
  • 适合语音合成与副语言事件检测研究者使用

副语言声音(如笑声、叹息)对生成更自然生动的语音至关重要。然而现有方法多依赖专有数据集,公开资源常存在语音不完整、时间戳不准或缺失、真实场景代表性不足等问题。为此,我们提出一种自动化框架,用于大规模生成副语言数据,并据此构建SynParaSpeech数据集。该数据集包含6类副语言内容,总计118.75小时音频,具有精确时间标记,全部源自自然对话语音。本工作首次实现大规模副语言数据的自动化构建,发布SynParaSpeech语料库,推动语音合成中副语言表达的自然化,提升语音理解中副语言事件检测性能。数据集及音频样例已开源:https://github.com/ShawnPi233/SynParaSpeech。

原文摘要 · Abstract (English)

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from incomplete speech, inaccurate or missing timestamps, and limited real-world relevance. To address these problems, we propose an automated framework for generating large-scale paralinguistic data and apply it to construct the SynParaSpeech dataset. The dataset comprises 6 paralinguistic categories with 118.75 hours of data and precise timestamps, all derived from natural conversational speech. Our contributions lie in introducing the first automated method for constructing large-scale paralinguistic datasets and releasing the SynParaSpeech corpus, which advances speech generation through more natural paralinguistic synthesis and enhances speech understanding by improving paralinguistic event detection. The dataset and audio samples are available at https://github.com/ShawnPi233/SynParaSpeech.

语音合成副语言数据集自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。