通过几何推理与多尺度融合,让模型更准识别注视目标物体。
Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

- 分两阶段建模,先对齐图像特征与语义实体,再融合多尺度信息
- 在四个数据集上AUC最高达0.987,参数量仅7.1M
- 适合关注视觉注意力机制与高效目标定位的研究者
注视目标估计旨在预测观察者在图像中注视的语义对象,该任务与人类注视的物体导向性密切相关。观察者倾向于选择特定语义实体作为注意目标,而非随机响应图像任意区域。然而,现有方法通常将此任务视为从全局特征到注视热图的直接映射,本质上将其当作像素级回归问题,未能显式表征被注视物体为独立实体,导致复杂场景下预测不稳定且语义不一致。为此,我们提出一种由物体语义引导的两阶段注视估计框架,将注视目标估计重构为层次化推理过程。方法在特征编码阶段引入物体级表示,使图像特征与离散语义实体对齐;随后结合多尺度特征融合及头姿与注视方向的几何约束,实现精细定位与物体级区分。在GazeFollow、VideoAttentionTarget、ChildPlay和GOO-Real上的实验表明,本方法分别取得0.961、0.948、0.987和0.977的AUC,所有基准表现优异,同时保持7.1M的紧凑参数量。
原文摘要 · Abstract (English)
Gaze target estimation aims to predict the semantic object an observer fixates upon within an image, a task deeply rooted in the object-oriented nature of human gaze. Observers tend to select a specific semantic entity as the attentional target, rather than responding randomly across arbitrary regions of the image. However, existing methods typically model this task as a direct mapping from global features to gaze heatmaps, essentially treating it as a pixel-level regression problem. This approach fails to explicitly represent the gazed object as a distinct entity, making it difficult to produce stable and semantically consistent predictions in complex scenes. To address this, we propose a two-stage gaze estimation framework guided by object semantics, reformulating gaze target estimation as a hierarchical reasoning process. Our method incorporates object-level representations during feature encoding to align image features with discrete semantic entities, then introduces multi-scale feature fusion and geometric constraints from head pose and gaze direction for fine-grained localization and object-level discrimination. Extensive experiments on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real demonstrate that our method achieves AUC of 0.961, 0.948, 0.987, and 0.977 respectively, delivering strong performance across all benchmarks while maintaining a compact parameter size of 7.1M.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。