arXiv:2505.10671cs.CV2025-05CVPR被引 5

通过3D场景上下文建模,实现远距离无眼图下的精准3D注视方向估计

GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding

  • 用3D姿态与物体位置构建场景上下文,学习主体与环境的空间关系
  • 在单帧下相比顶尖方法平均角度误差降低13%~37%
  • 适合无近距离眼图、复杂视角的现实场景注视估计任务

我们提出一种新的3D注视方向估计方法,通过学习主体与场景中物体之间的空间关系来输出3D注视方向。该方法针对非受限场景,包括主体远离或背对摄像头时无法获取近距眼部图像的情况。以往方法通常仅依赖2D外观特征,或在不可学习的后处理阶段引入有限的空间线索(如深度图)。在这些情况下,从2D观测推断3D注视方向极具挑战性:主体姿态、场景布局、注视方向及相机姿态的变化,导致相同3D场景产生多样化的2D外观和3D注视方向。为此,我们提出GA3CE(Gaze-Aware 3D Context Encoding):将主体与场景表示为3D姿态和物体位置,作为3D上下文以学习3D空间中的关系。受人类视觉启发,我们将此上下文对齐至主观视角空间,显著降低空间复杂度。此外,我们提出方向-距离解耦(D$^3$)位置编码,更有效地捕捉3D上下文与注视方向在方向与距离空间中的关系。实验表明,在基准数据集上,单帧设置下相较领先基线方法,平均角度误差降低13%~37%。

原文摘要 · Abstract (English)

We propose a novel 3D gaze estimation approach that learns spatial relationships between the subject and objects in the scene, and outputs 3D gaze direction. Our method targets unconstrained settings, including cases where close-up views of the subject's eyes are unavailable, such as when the subject is distant or facing away. Previous approaches typically rely on either 2D appearance alone or incorporate limited spatial cues using depth maps in the non-learnable post-processing step. Estimating 3D gaze direction from 2D observations in these scenarios is challenging; variations in subject pose, scene layout, and gaze direction, combined with differing camera poses, yield diverse 2D appearances and 3D gaze directions even when targeting the same 3D scene. To address this issue, we propose GA3CE: Gaze-Aware 3D Context Encoding. Our method represents subject and scene using 3D poses and object positions, treating them as 3D context to learn spatial relationships in 3D space. Inspired by human vision, we align this context in an egocentric space, significantly reducing spatial complexity. Furthermore, we propose D$^3$ (direction-distance-decomposed) positional encoding to better capture the spatial relationship between 3D context and gaze direction in direction and distance space. Experiments demonstrate substantial improvements, reducing mean angle error by 13%-37% compared to leading baselines on benchmark datasets in single-frame settings.

3D注视估计场景上下文姿态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。