arXiv:2604.02397cs.CVcs.AI2026-04

VE-MD通过结构化编码避免个体识别,实现隐私友好的群体情绪识别。

Variational Encoder--Multi-Decoder (VE-MD) for Privacy-by-functional-design (Group) Emotion Recognition

  • 用共享潜空间联合学习情绪与身体面部结构,仅输出群体情绪
  • 在GAF-3.0上达90.06%,多模态融合提升至82.25%
  • 适合需保护隐私的公共场景情绪分析

群体情绪识别(GER)旨在推断教室、人群和公共活动等社交环境中的集体情感。现有方法多依赖个体层面处理,如裁剪人脸、人物追踪或逐人特征提取,导致分析流程以个人为中心,在仅需群体理解的场景中引发隐私担忧。本文提出变分编码器-多解码器(VE-MD)框架,基于隐私感知的功能设计实现群体情绪识别。该模型不提供正式匿名化或密码学隐私保障,而是通过约束仅预测聚合群体情绪、不进行身份识别或个体情绪输出,避免显式个体监控。VE-MD学习共享潜表示,联合优化情绪分类与内部身体及面部结构预测。研究两种结构解码策略:基于Transformer的PersonQuery解码器与可适应不同群体规模的密集热图解码器。在六个真实场景数据集(含两个GER和四个个体情绪识别IER基准)上的实验表明,结构监督显著提升表示学习效果。更重要的是,揭示了GER与IER的关键差异:仅优化潜空间常不足以支持GER,因其易削弱交互相关线索;而保留显式结构输出能有效提升集体情感推断。相反,投影后的结构表示对IER具有有效去噪作用。VE-MD在GAF-3.0上达到90.06%的SOTA性能,多模态融合下在VGAF上达82.25%。结果表明,保留交互相关结构信息对群体情感建模至关重要,且无需依赖个体特征提取。在多模态融合音频的IER数据集上,VE-MD在SamSemo上达77.9%(加文本模态),在MER-MULTI(63.8%)、DFEW(70.7%)和EngageNet(69.0%)上表现竞争力。

原文摘要 · Abstract (English)

Group Emotion Recognition (GER) aims to infer collective affect in social environments such as classrooms, crowds, and public events. Many existing approaches rely on explicit individual-level processing, including cropped faces, person tracking, or per-person feature extraction, which makes the analysis pipeline person-centric and raises privacy concerns in deployment scenarios where only group-level understanding is needed. This research proposes VE-MD, a Variational Encoder-Multi-Decoder framework for group emotion recognition under a privacy-aware functional design. Rather than providing formal anonymization or cryptographic privacy guarantees, VE-MD is designed to avoid explicit individual monitoring by constraining the model to predict only aggregate group-level affect, without identity recognition or per-person emotion outputs. VE-MD learns a shared latent representation jointly optimized for emotion classification and internal prediction of body and facial structural representations. Two structural decoding strategies are investigated: a transformer-based PersonQuery decoder and a dense Heatmap decoder that naturally accommodates variable group sizes. Experiments on six in-the-wild datasets, including two GER and four Individual Emotion Recognition (IER) benchmarks, show that structural supervision consistently improves representation learning. More importantly, the results reveal a clear distinction between GER and IER: optimizing the latent space alone is often insufficient for GER because it tends to attenuate interaction-related cues, whereas preserving explicit structural outputs improves collective affect inference. In contrast, projected structural representations seem to act as an effective denoising bottleneck for IER. VE-MD achieves state-of-the-art performance on GAF-3.0 (up to 90.06%) and VGAF (82.25% with multimodal fusion with audio). These results show that preserving interaction-related structural information is particularly beneficial for group-level affect modeling without relying on prior individual feature extraction. On IER datasets using multimodal fusion with audio modality, VE-MD outperforms SOTA on SamSemo (77.9%, adding text modality) while achieving competitive performances on MER-MULTI (63.8%), DFEW (70.7%) and EngageNet (69.0).

群体情绪识别隐私保护结构化表示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。