arXiv:2602.21464eess.AScs.CL2026-02中稿 · Speech Prosody 202…被引 1

首个基于真实比赛结果的自发情绪语音数据集,支持多模态情感分析。

iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis

  • 基于真实赛事结果采集自发语音,避免表演式情绪干扰。
  • 包含语音转录、角色分离与词级对齐等丰富元数据。
  • 可同步匹配微表情标注,适合研究语音与手势的情感动态。

本文提出 iMiGUE-Speech,是 iMiGUE 数据集的扩展版本,构建了一个用于研究情感与情绪状态的自发性情感语料库。新版本聚焦语音数据,补充了语音转录、采访者与受访者的角色分离以及词级强制对齐等额外元数据。与依赖表演或实验室诱发情绪的现有语音数据集不同,iMiGUE-Speech 捕捉的是由真实比赛结果自然引发的自发情绪。为验证数据集价值并建立初步基准,本文引入两项评估任务:语音情感识别和基于文本的语气分析。这两项任务利用前沿预训练表示方法,评估数据集在声学与语言模态中捕捉自发情感状态的能力。iMiGUE-Speech 可与原始 iMiGUE 数据集中的微表情标注同步配对,形成独特的多模态资源,用于研究语音-手势情感动态。扩展数据集已公开于 https://github.com/CV-AC/imigue-speech。

原文摘要 · Abstract (English)

This work presents iMiGUE-Speech, an extension of the iMiGUE dataset that provides a spontaneous affective corpus for studying emotional and affective states. The new release focuses on speech and enriches the original dataset with additional metadata, including speech transcripts, speaker-role separation between interviewer and interviewee, and word-level forced alignments. Unlike existing emotional speech datasets that rely on acted or laboratory-elicited emotions, iMiGUE-Speech captures spontaneous affect arising naturally from real match outcomes. To demonstrate the utility of the dataset and establish initial benchmarks, we introduce two evaluation tasks for comparative assessment: speech emotion recognition and transcript-based sentiment analysis. These tasks leverage state-of-the-art pre-trained representations to assess the dataset's ability to capture spontaneous affective states from both acoustic and linguistic modalities. iMiGUE-Speech can also be synchronously paired with micro-gesture annotations from the original iMiGUE dataset, forming a uniquely multimodal resource for studying speech-gesture affective dynamics. The extended dataset is available at https://github.com/CV-AC/imigue-speech.

情感分析语音数据集多模态自发情绪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。