arXiv:2607.00296cs.CVcs.AI2026-07

提出门控机制,让模型学会在合适时机融合面部情绪信息提升动作预测

Learning When to Listen: Gated Affect Fusion for Human Motion Prediction

论文配图:Learning When to Listen: Gated Affect Fusion for Human Motion Prediction
图 1 · 摘自论文原文
  • 设计门控情感变换器,动态调节情绪与姿态信息的融合时机
  • 实验表明情绪信息仅在30帧内有效,过长则反而降低精度
  • 适合做复杂场景下人体动作预测的多模态研究者参考

在非受限真实视频中进行人体运动预测仍具挑战性,源于未来行为的不确定性及多模态观测中的噪声。尽管面部情绪可能提供补充行为线索,但其在运动预测框架中的实际效用与时间边界尚不明确。本文系统研究了野生环境下情绪条件化预测的实用性与时序限制。构建结合MediaPipe体姿轨迹与HSEmotion面部情绪表示的多模态管道,引入门控情感变换器(GAT),动态调控跨模态信息流。在严格的受试者独立协议下进行多时域评估发现,简单的早期跨模态拼接会持续降低预测精度,而我们的门控机制通过自适应控制情绪流实现稳定融合。关键的反事实实验显示,当情绪输入被打乱时,学习到的门控能有效抑制非结构化噪声,同时对合理的情绪信号保持响应。实证结果表明,面部情绪特征仅在短至中等时间窗口(如30帧)内提供有限的预测线索,长期轨迹仍主要由内在运动连续性主导。研究证明,面部情绪应被视为互补行为线索而非主导因素,为非受限场景下选择性多模态融合提供了实践指导。

原文摘要 · Abstract (English)

Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence of noisy multimodal observations. While facial affect potentially provides complementary behavioral cues, its practical utility and mechanistic boundaries within motion forecasting frameworks remain poorly understood. In this work, we present a systematic study investigating the utility and temporal limitations of affect-conditioned forecasting in-the-wild. We establish a rigorous multimodal pipeline combining MediaPipe body pose trajectories with HSEmotion facial affect representations, and introduce the Gated Affect Transformer (GAT) to dynamically regulate cross-modal information flow. Through extensive multi-horizon evaluations under a strict subject-wise protocol, we demonstrate that naive early cross-modal concatenation consistently degrades forecasting accuracy relative to pose-only baselines. Conversely, our proposed gating mechanism stabilizes cross-modal integration by adaptively controlling the affective stream. Crucially, controlled counterfactual experiments using shuffled and randomized affect inputs reveal that the learned gate successfully suppresses unstructured cross-modal noise while remaining responsive to plausible affective signals. Furthermore, our empirical results indicate that facial affect features provide bounded, horizon-dependent predictive cues strictly within short-to-medium windows (e.g., 30 frames), whereas long-term trajectories remain predominantly governed by intrinsic kinematic continuity. Our findings provide empirical evidence that facial affect should be regarded as a complementary behavioral cue rather than a dominant driver of future motion, offering practical guidance for selective multimodal fusion in unconstrained human motion forecasting.

动作预测多模态融合情绪识别门控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。