arXiv:2511.00256eess.AScs.LG2025-11被引 6

首个大规模自然播客数据集,助力情感化语音转换研究

NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion

  • 构建5049小时真实播客数据,自动标注情绪与语音质量
  • 支持训练更自然、有表现力的语音转换模型
  • 适合语音转换、情感建模及真实场景研究者使用

日常语音不仅传递语义,还体现说话人身份、情感状态与交流情境。然而,现有语音数据集多为表演性、规模有限,难以捕捉真实交流中的表达丰富性。尽管大尺寸神经网络推动了大规模语音语料的发展,但语音转换(VC)领域仍缺乏大规模、富有表现力且贴近现实的资源来建模自然语调与情感。为此,我们发布NaturalVoices(NV),首个专为情感感知语音转换设计的大规模自发播客数据集。该数据集包含5049小时自发录制的播客内容,涵盖数千名说话人、多样话题与自然语用风格,提供情绪(类别与属性)、语音质量、转录文本、说话人身份与声学事件的自动标注。我们还开放了模块化标注工具与灵活筛选管道,支持研究人员定制子集用于多种VC任务。实验表明,NaturalVoices可支持鲁棒且泛化的语音转换模型,生成自然、富有表现力的语音;同时揭示当前架构在大规模自发数据上的局限性。结果表明,NaturalVoices既是宝贵资源,也是推进语音转换领域的挑战性基准。数据集已公开:https://huggingface.co/JHU-SmileLab

原文摘要 · Abstract (English)

Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab

语音转换情感建模播客数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。