arXiv:2603.22536eess.AScs.SD2026-03

构建70小时对话音频数据集,实现情绪的连续动态标注。

MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition

  • 采用连续时间标注捕捉情绪的动态变化过程。
  • 包含超过70小时自然对话音频,支持上下文相关的情绪分析。
  • 适合研究真实场景下情绪识别的学者与开发者使用。

情感计算旨在理解与建模人类情绪以服务于计算系统。在该领域中,语音情绪识别(SER)专注于预测通过语音传达的情绪。早期的SER系统依赖于有限的数据集和传统机器学习模型,而近期的深度学习方法则需要大规模、自然情境下的情绪语料库。为满足这一需求,本文引入MSP-Conversation语料库:一个超过70小时的对话音频数据集,配有时间连续的情绪标注与详细的说话人分离信息。时间连续标注捕捉了情绪表达的动态性与情境依赖性,标注内容包括愉悦度、唤醒度和支配度的细粒度时间轨迹。音频数据来源于公开播客,且与MSP-Podcast语料库中的独立发言片段有重叠,便于直接比较上下文与非上下文标注方法的效果。本文详细描述了语料库的构建过程、标注方法、标注分析及基线SER实验,确立了MSP-Conversation语料库在自然情境下动态情绪识别研究中的重要价值。

原文摘要 · Abstract (English)

Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited datasets and traditional machine learning models, recent deep learning approaches demand largescale, naturalistic emotional corpora. To address this need, we introduce the MSP-Conversation corpus: a dataset of more than 70 hours of conversational audio with time-continuous emotional annotations and detailed speaker diarizations. The time-continuous annotations capture the dynamic and contextdependent nature of emotional expression. The annotations in the corpus include fine-grained temporal traces of valence, arousal, and dominance. The audio data is sourced from publicly available podcasts and overlaps with a subset of the isolated speaking turns in the MSP-Podcast corpus to facilitate direct comparisons between annotation methods (i.e., in-context versus out-of-context annotations). The paper outlines the development of the corpus, annotation methodology, analyses of the annotations, and baseline SER experiments, establishing the MSP-Conversation corpus as a valuable resource for advancing research in dynamic SER in naturalistic settings.

情绪识别语音分析自然对话数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。