arXiv:2601.13225cs.CVcs.HC2026-01中稿 · publication at IEE…被引 6

构建首个标注情绪混合强度的多模态数据集,推动复杂情绪识别研究

Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations

  • 提出多模态情绪混合数据集,含10种混合情绪及三种强度比例标注
  • 模型在情绪存在性预测上最高达35%准确率,强度判断达18%准确率
  • 适合情绪计算、人机交互领域研究者使用,助力真实场景情感理解

人类常同时体验多种情绪的混合状态,且各情绪强度不同。然而,现有视频情绪识别方法多仅支持单一情绪识别,少数尝试混合情绪的方法也缺乏对情绪强度相对关系的标注。为弥补这一不足,本文构建了BLEMORE数据集,包含58位演员表演的3000余段视频音频片段,涵盖6种基本情绪与10种混合情绪组合,每种混合情绪有50/50、70/30、30/70三种强度配置。我们在此数据集上评估主流视频分类模型在两项任务上的表现:(1) 预测情绪是否存在;(2) 预测情绪在混合中的相对强度。结果显示,单模态模型在验证集上情绪存在性准确率最高达29%,强度判断达13%;多模态方法显著提升,ImageBind + WavLM 达35%存在性准确率,HiCMAE 达18%强度准确率。在测试集上,最优模型取得33%存在性准确率(VideoMAEv2 + HuBERT)和18%强度准确率(HiCMAE)。BLEMORE为复杂情绪识别研究提供了重要资源。

原文摘要 · Abstract (English)

Humans often experience not just a single basic emotion at a time, but rather a blend of several emotions with varying salience. Despite the importance of such blended emotions, most video-based emotion recognition approaches are designed to recognize single emotions only. The few approaches that have attempted to recognize blended emotions typically cannot assess the relative salience of the emotions within a blend. This limitation largely stems from the lack of datasets containing a substantial number of blended emotion samples annotated with relative salience. To address this shortcoming, we introduce BLEMORE, a novel dataset for multimodal (video, audio) blended emotion recognition that includes information on the relative salience of each emotion within a blend. BLEMORE comprises over 3,000 clips from 58 actors, performing 6 basic emotions and 10 distinct blends, where each blend has 3 different salience configurations (50/50, 70/30, and 30/70). Using this dataset, we conduct extensive evaluations of state-of-the-art video classification approaches on two blended emotion prediction tasks: (1) predicting the presence of emotions in a given sample, and (2) predicting the relative salience of emotions in a blend. Our results show that unimodal classifiers achieve up to 29% presence accuracy and 13% salience accuracy on the validation set, while multimodal methods yield clear improvements, with ImageBind + WavLM reaching 35% presence accuracy and HiCMAE 18% salience accuracy. On the held-out test set, the best models achieve 33% presence accuracy (VideoMAEv2 + HuBERT) and 18% salience accuracy (HiCMAE). In sum, the BLEMORE dataset provides a valuable resource to advancing research on emotion recognition systems that account for the complexity and significance of blended emotion expressions.

情绪识别多模态数据集混合情绪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。