arXiv:2606.07585cs.CVcs.AI2026-06

不依赖个人特征,用集体音视频信号实现隐私安全的群体情绪识别

Multimodal Group Emotion Recognition In-the-Wild Towards a Privacy-Safe Non-Individual Approach

论文配图:Multimodal Group Emotion Recognition In-the-Wild Towards a Privacy-Safe Non-Individual Approach
图 1 · 摘自论文原文
  • 用跨注意力融合音视频信号,结合帧注意力池化处理时间信息
  • 在真实场景下实现高鲁棒性群体情绪识别,无需个体特征输入
  • 适合关注隐私保护的智能监控、公共空间情绪分析等场景

本论文聚焦于真实场景下的群体情绪识别(GER),强调隐私保护。与依赖人脸、注视或语音等个体线索的传统方法不同,本文利用集体音视频信号推断群体情绪,降低个体监控风险。提出两种互补框架:第一是基于跨注意力的多模态融合架构,结合帧注意力池化(FAP)进行时序聚合,通过合成数据增强和消融实验验证其在真实环境中的鲁棒性;第二是变分编码器多解码器(VE-MD)框架,学习共享潜在空间以同时完成情绪分类与结构表征预测(包括身体和面部线索),探索基于DETR和热图的两种解码策略,分析结构表征在群体与个体场景中的作用。主要贡献包括:阐明多模态与结构线索在群体情感计算中的作用;提出两种隐私友好的多模态GER架构;证明在不使用个体特征作为输入的前提下仍可实现竞争力表现。

原文摘要 · Abstract (English)

This thesis addresses group emotion recognition (GER) in-the-wild with a focus on privacy preservation. Unlike traditional emotion recognition methods that rely on individual-level cues such as face, gaze, or voice analysis, this work uses collective audio-video signals to infer emotions at the group level, reducing risks of individual monitoring and surveillance. Two complementary frameworks are proposed. The first is a cross-attention multimodal architecture for audio-video fusion, combined with Frames Attention Pooling (FAP) for temporal aggregation. It is supported by synthetic data augmentation and validated through ablation studies, demonstrating robustness in real-world GER conditions. The second framework, Variational Encoder Multi-Decoder (VE-MD), learns a shared latent space for emotion classification and structural representation prediction, including body and face cues. Two decoding strategies, DETR-based and heatmap-based, are explored to analyze the role of structural representations in group and individual settings. The thesis makes three main contributions: it clarifies the role of multimodality and structural cues in group-level affective computing; introduces two architectures for privacy-preserving multimodal GER; and shows that competitive performance can be achieved without using individual features as input data.

群体情绪识别多模态融合隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。