不依赖个人特征,用集体音视频信号实现隐私安全的群体情绪识别
Multimodal Group Emotion Recognition In-the-Wild Towards a Privacy-Safe Non-Individual Approach

- 用跨注意力融合音视频信号,结合帧注意力池化处理时间信息
- 在真实场景下实现高鲁棒性群体情绪识别,无需个体特征输入
- 适合关注隐私保护的智能监控、公共空间情绪分析等场景
本论文聚焦于真实场景下的群体情绪识别(GER),强调隐私保护。与依赖人脸、注视或语音等个体线索的传统方法不同,本文利用集体音视频信号推断群体情绪,降低个体监控风险。提出两种互补框架:第一是基于跨注意力的多模态融合架构,结合帧注意力池化(FAP)进行时序聚合,通过合成数据增强和消融实验验证其在真实环境中的鲁棒性;第二是变分编码器多解码器(VE-MD)框架,学习共享潜在空间以同时完成情绪分类与结构表征预测(包括身体和面部线索),探索基于DETR和热图的两种解码策略,分析结构表征在群体与个体场景中的作用。主要贡献包括:阐明多模态与结构线索在群体情感计算中的作用;提出两种隐私友好的多模态GER架构;证明在不使用个体特征作为输入的前提下仍可实现竞争力表现。
原文摘要 · Abstract (English)
This thesis addresses group emotion recognition (GER) in-the-wild with a focus on privacy preservation. Unlike traditional emotion recognition methods that rely on individual-level cues such as face, gaze, or voice analysis, this work uses collective audio-video signals to infer emotions at the group level, reducing risks of individual monitoring and surveillance. Two complementary frameworks are proposed. The first is a cross-attention multimodal architecture for audio-video fusion, combined with Frames Attention Pooling (FAP) for temporal aggregation. It is supported by synthetic data augmentation and validated through ablation studies, demonstrating robustness in real-world GER conditions. The second framework, Variational Encoder Multi-Decoder (VE-MD), learns a shared latent space for emotion classification and structural representation prediction, including body and face cues. Two decoding strategies, DETR-based and heatmap-based, are explored to analyze the role of structural representations in group and individual settings. The thesis makes three main contributions: it clarifies the role of multimodality and structural cues in group-level affective computing; introduces two architectures for privacy-preserving multimodal GER; and shows that competitive performance can be achieved without using individual features as input data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。