构建大规模视频情感数据集,助力群体情绪智能识别。
GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos

- 融合视频、音频与上下文信息,标注群体情绪三元极性与离散情绪。
- 提出CAGNet模型,在GAViD上达63.20%测试准确率。
- 适合研究社交计算、多模态情感分析的学者使用。
理解真实社会系统中的情感动态是建模和分析复杂环境中人与人交互的基础。群体情感源于交织的人际互动、情境影响与行为线索,其量化建模属于具有挑战性的计算社会系统问题。然而,由于缺乏大规模标注数据集及多模态社交互动中情境与行为变异性带来的复杂性,真实场景下的群体情感计算建模仍面临困难。现有数据集在多模态与情境信息标注方面不完整,制约了领域进展。为此,我们推出群组情感视频数据集(GAViD),包含5091个视频片段,涵盖视频、音频与情境信息,标注有三元极性与离散情绪标签,并通过VideoGPT生成情境元数据及人工标注动作线索进行增强。同时提出上下文感知群组情感识别网络(CAGNet),在GAViD上实现63.20%测试准确率,达到当前最优水平。数据集与代码已开源至github.com/deepakkumar-iitr/GAViD。
原文摘要 · Abstract (English)
Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, contextual influences, and behavioral cues, making its quantitative modeling a challenging computational social systems problem. However, computational modeling of group affect in in-the-wild scenarios remains challenging due to limited large-scale annotated datasets and the inherent complexity of multimodal social interactions shaped by contextual and behavioral variability. The lack of comprehensive datasets annotated with multimodal and contextual information further limits advances in the field. To address this, we introduce the Group Affect from ViDeos (GAViD) dataset, comprising 5091 video clips with multimodal data (video, audio and context), annotated with ternary valence and discrete emotion labels and enriched with VideoGPT-generated contextual metadata and human-annotated action cues. We also present Context-Aware Group Affect Recognition Network (CAGNet) for multimodal context-aware group affect recognition. CAGNet achieves 63.20\% test accuracy on GAViD, comparable to state-of-the-art performance. The dataset and code are available at github.com/deepakkumar-iitr/GAViD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。