arXiv:2608.02331cs.CVcs.AI2026-08

通过空间上下文增强身体情绪识别,提升静态图像下的情感判别能力

Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild

论文配图:Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild
图 1 · 摘自论文原文
  • 引入多领域专家混合模型,用场景与物体生成情绪分布先验
  • 在Body Language Database上达到0.3269的情绪识别得分
  • 适合关注静态图像中情感分析的视觉研究者

同一身体姿态在不同上下文中可能表达完全不同的情绪,但现有身体情绪识别方法通常将场景与物体线索视为辅助特征,而非结构化先验。本文提出上下文感知的领域专家混合模型(CA-MoDE),通过专门的场景与物体专家生成基于各自领域的软情绪分布。这些领域条件化的软预测作为结构化上下文先验,在分布层面而非特征层面调节身体专家的输出。为融合多领域信号,我们设计任务定制的极大支持门控策略,针对每个情绪维度选择最强的上下文信号,避免冲突或无信息上下文分布平均带来的信号稀释。在Body Language Database上,CA-MoDE取得0.3269的情绪识别得分,仅使用单张静态图像即超越现有时序模型,表明显式建模结构化空间上下文可作为行为动态的互补判别代理。

原文摘要 · Abstract (English)

The same body posture can convey entirely different emotions depending on its surrounding context, yet most methods for recognising bodily emotions treat scene and object cues as auxiliary feature augmentations rather than as structured priors over the plausibility of emotions. We introduce the Context-Aware Mixture of Domain Experts (CA-MoDE) for bodily emotion recognition. CA-MoDE incorporates dedicated scene and object experts to generate soft distributions over emotion categories conditioned on their respective domains. These domain-conditioned soft predictions serve as structured contextual priors that modulate the body expert's predictions at the distributional level rather than at the feature level. To fuse these multi-domain signals, we propose a task-tailored max-endorsement gating strategy that selects the strongest contextual signal across experts for each emotion dimension. Our gating strategy mitigates the signal dilution that typically occurs when conflicting or uninformative context distributions are averaged. CA-MoDE achieves an Emotion Recognition Score of 0.3269 on the Body Language Database. By outperforming existing temporal models using only single still images, our framework demonstrates that explicitly modelling structured spatial context can serve as a complementary discriminative proxy for the behavioural dynamics typically captured by video.

情绪识别上下文建模静态图像多专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。