arXiv:2511.05955cs.CVcs.LG2025-11

融合场景与人脸信息,提升多人对话中社交凝视预测准确率

CSGaze: Context-aware Social Gaze Prediction

  • 结合人脸、场景与上下文线索,建模多人对话中的凝视动态
  • 在GP-Static、UCO-LAEO等数据集上达到领先性能
  • 通过注意力热力图揭示模型决策逻辑,适合人机交互研究

人的凝视能反映其注意力焦点、社交参与度和自信程度。本文研究如何结合上下文线索、视觉场景和面部信息,有效预测与解释对话中的社交凝视模式。我们提出CSGaze,一种基于上下文感知的多模态方法,利用人脸与场景信息作为互补输入,提升从多人图像中预测社交凝视的能力。模型引入以主讲者为中心的细粒度注意力机制,更好捕捉社交凝视动态。实验表明,CSGaze在GP-Static、UCO-LAEO和AVA-LAEO数据集上表现优于或媲美现有最优方法。研究验证了上下文线索对提升社交凝视预测的关键作用。此外,通过生成注意力分数实现初步可解释性,揭示模型决策过程。我们在开放集数据集上测试模型泛化能力,证明其在多样化场景下的鲁棒性。

原文摘要 · Abstract (English)

A person's gaze offers valuable insights into their focus of attention, level of social engagement, and confidence. In this work, we investigate how contextual cues combined with visual scene and facial information can be effectively utilized to predict and interpret social gaze patterns during conversational interactions. We introduce CSGaze, a context aware multimodal approach that leverages facial, scene information as complementary inputs to enhance social gaze pattern prediction from multi-person images. The model also incorporates a fine-grained attention mechanism centered on the principal speaker, which helps in better modeling social gaze dynamics. Experimental results show that CSGaze performs competitively with state-of-the-art methods on GP-Static, UCO-LAEO and AVA-LAEO. Our findings highlight the role of contextual cues in improving social gaze prediction. Additionally, we provide initial explainability through generated attention scores, offering insights into the model's decision-making process. We also demonstrate our model's generalizability by testing our model on open set datasets that demonstrating its robustness across diverse scenarios.

凝视预测多模态注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。