arXiv:2607.22847cs.CV2026-07中稿 · ICPR

通过联合建模视线与隐含关系,从静态图像中解码社会意图。

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

论文配图:Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling
图 1 · 摘自论文原文
  • 以目标为中心的框架,用关系注意力捕捉人与人之间的细微关联。
  • 在扩展数据集上达到顶尖性能,首次量化证明可从视线模式中学习社会层级。
  • 适合研究视觉社交推理、多主体交互建模的学者参考。

人类视线不仅指向视觉目标,更是在静态图像中微妙地体现社会意图,而现有模型通常独立处理个体,将视线视为独立同分布的量或孤立预测社会语义。近期多主体方法虽尝试解决此问题,但常将社会关系视为僵化的后处理分类,与视线估计过程脱钩。这种简化忽略了社会意图作为视线行为潜在驱动力的本质。为此,我们提出 ANCHOR,一种以目标为中心的范式,通过联合建模视觉注意与潜在隐含关系来解码视线锚定的社会意图。该方法将这些依赖关系揭示为视线行为的潜在结构骨架。架构采用关系注意力机制捕捉细粒度人际联系,并通过特征调制实现单个视觉主干网络的高效多主体解析。为稳定耦合训练,引入优化协同机制,解决空间视线精度与潜在社会推理间的内在冲突。该方法通过寻找稳定平坦的极小值,同时协调竞争任务梯度,确保鲁棒泛化。我们在一个包含密集多人标注和新型社会影响力排序的扩展基准上验证了框架有效性。结果表明,本方法达到最先进性能,并首次提供定量证据,证明隐含社会层级可从静态视线模式中稳健解耦并学习。

原文摘要 · Abstract (English)

Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals independently, treating gaze as an i.i.d. quantity or predicting social semantics in isolation. Recent multi-person methods attempt to address this but often treat social relations as rigid, post-hoc classifications decoupled from the gaze estimation process. This oversimplification fails to capture the nuanced nature of social intent, which acts as an underlying driver of gaze behavior rather than a secondary categorical output. We address these limitations by proposing ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations. Our approach surfaces these dependencies as the latent structural scaffolding of gaze behavior. The architecture utilizes a relational attention mechanism to capture fine-grained interpersonal links, leveraging feature-wise modulation for efficient multi-person parsing from a single vision backbone. To stabilize the training of this coupled formulation, we implement an optimization synergy to resolve the inherent conflicts between spatial gaze accuracy and latent social reasoning. This approach ensures robust generalization by seeking stable, flat minima while simultaneously harmonizing competing task gradients. We validate our framework on an extended benchmark featuring dense multi-person annotations and novel social influence rankings. Our results demonstrate state-of-the-art performance and provide the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.

社会意图视线分析关系建模多主体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。