arXiv:2604.04160eess.AScs.SD2026-04被引 1

构建大规模情感语音数据集,支持细粒度情绪描述与合成。

AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis

  • 采用人机协同标注流程,生成六维细粒度情绪描述。
  • 模型在情绪描述与合成任务上显著优于现有方法。
  • 适合研究语音情感建模、生成与自然语言交互的学者。

情绪在口语交流中至关重要,但现有语音情感建模多依赖预定义类别或低维连续属性,表达能力有限。近年来,语音情感描述与合成研究发现,自然语言描述能更灵活、可解释地表征语音情感特征。然而,这一方向受限于缺乏高质量、细粒度的自然语言标注情感语音数据集。为此,我们推出AffectSpeech,一个大规模真人录制语音语料库,配备结构化细粒度情感描述,支持多粒度语音表达建模。每个语音片段涵盖六个互补维度:情感极性、开放式词汇情绪标签、强度等级、韵律属性、显著段落和语义内容。为平衡标注质量与可扩展性,我们设计了人机协同标注流程,结合算法预标注、多大模型生成描述及人工验证。此外,标注内容被重构成多样描述风格,提升语言多样性并降低下游建模中的风格偏差。在语音情感描述与合成任务上的实验表明,基于AffectSpeech训练的模型在多种评估设置下均表现优异。

原文摘要 · Abstract (English)

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech emotion captioning and synthesis have shown that textual descriptions provide a more flexible and interpretable alternative for representing affective characteristics in speech. However, progress in this direction is hindered by the lack of an emotional speech dataset aligned with reliable and fine-grained natural language annotations. To tackle this, we introduce AffectSpeech, a large-scale corpus of human-recorded speech enriched with structured descriptions for fine-grained emotion analysis and generation. Each utterance is characterized across six complementary dimensions, including sentiment polarity, open-vocabulary emotion captions, intensity level, prosodic attributes, prominent segments, and semantic content, enabling multi-granular modeling of vocal expression. To balance annotation quality and scalability, we adopt a human-LLM collaborative annotation pipeline that integrates algorithmic pre-labeling, multi-LLM description generation, and human-in-the-loop verification. Furthermore, these annotations are reformulated into diverse descriptive styles to enhance linguistic diversity and reduce stylistic bias in downstream modeling. Experimental results on speech emotion captioning and synthesis demonstrate that models trained on AffectSpeech consistently achieve superior performance across multiple evaluation settings.

语音情感数据集自然语言描述生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。