多模态模型联合预测个人到事件的情绪,提升社交场景理解
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- 用多模态Swin Transformer与S3Attention融合场景、人物和上下文信息
- 在VGAF-GEMS数据集上实现个体、群体和事件级情绪的精准预测
- 适合研究社交智能、人机交互及情感计算的学者与开发者
理解个体、群体和事件层面的情绪及其上下文信息,对分析多人社交情境至关重要。本文将情绪理解任务定义为从细粒度个体情绪到粗粒度群体与事件情绪的预测。提出GEMS框架,采用基于多模态Swin-Transformer和S3Attention的架构,联合处理输入场景、群体成员及上下文信息,生成联合预测结果。现有针对多人情绪的基准数据集主要关注原子交互行为,侧重于时间维度上的情绪感知与群体层面标注。为此,本文扩展并提出VGAF-GEMS,在原有VGAF数据集群体标注基础上,提供更细粒度且全面的分析。GEMS旨在预测基本离散情绪与连续情绪(包括效价与唤醒度),以及个体、群体和事件层面的感知情绪。通过定量与定性对比,验证了该框架在VGAF-GEMS基准上的有效性。代码与数据已公开。
原文摘要 · Abstract (English)
Understanding individual, group and event level emotions along with contextual information is crucial for analyzing a multi-person social situation. To achieve this, we frame emotion comprehension as the task of predicting fine-grained individual emotion to coarse grained group and event level emotion. We introduce GEMS that leverages a multimodal swin-transformer and S3Attention based architecture, which processes an input scene, group members, and context information to generate joint predictions. Existing multi-person emotion related benchmarks mainly focus on atomic interactions primarily based on emotion perception over time and group level. To this end, we extend and propose VGAF-GEMS to provide more fine grained and holistic analysis on top of existing group level annotation of VGAF dataset. GEMS aims to predict basic discrete and continuous emotions (including valence and arousal) as well as individual, group and event level perceived emotions. Our benchmarking effort links individual, group and situational emotional responses holistically. The quantitative and qualitative comparisons with adapted state-of-the-art models demonstrate the effectiveness of GEMS framework on VGAF-GEMS benchmarking. We believe that it will pave the way of further research. The code and data is available at: https://github.com/katariaak579/GEMS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。