arXiv:2502.20249cs.CV2025-02CVPR被引 17

用弱监督提升真实场景下的3D视线估计,效果优于现有方法。

Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labels

  • 通过自训练框架利用2D注视追踪数据生成3D伪标签。
  • 在Gaze360和GFIE上实现当前最优的跨域泛化性能。
  • 提出统一架构,同时学习图像与视频中的动态静态视线信息。

真实世界中准确进行3D视线估计仍面临外观差异、头部姿态变化、遮挡及缺乏真实场景3D视线数据集等挑战。为此,我们提出一种新型自训练弱监督视线估计框架(ST-WSGE),该两阶段框架利用多样化的2D视线数据(如注视跟随数据),包含丰富的外观、自然场景与视线分布,并提出生成3D伪标签的方法以增强模型泛化能力。此外,传统针对图像或视频分别设计的模态专用模型限制了训练数据的有效利用。为此,我们提出视线变换器(GaT),一种模态无关架构,可同时从图像与视频数据中学习静态与动态视线信息。结合3D视频数据集与来自注视跟随任务的2D视线目标标签,本方法在无约束基准测试(如Gaze360与GFIE)上实现了显著的先进水平性能,尤其在视频视线估计中取得明显跨模态提升;在MPIIFaceGaze与Gaze360等数据集上也优于前视人脸方法。代码与预训练模型将向社区开放。

原文摘要 · Abstract (English)

Accurate 3D gaze estimation in unconstrained real-world environments remains a significant challenge due to variations in appearance, head pose, occlusion, and the limited availability of in-the-wild 3D gaze datasets. To address these challenges, we introduce a novel Self-Training Weakly-Supervised Gaze Estimation framework (ST-WSGE). This two-stage learning framework leverages diverse 2D gaze datasets, such as gaze-following data, which offer rich variations in appearances, natural scenes, and gaze distributions, and proposes an approach to generate 3D pseudo-labels and enhance model generalization. Furthermore, traditional modality-specific models, designed separately for images or videos, limit the effective use of available training data. To overcome this, we propose the Gaze Transformer (GaT), a modality-agnostic architecture capable of simultaneously learning static and dynamic gaze information from both image and video datasets. By combining 3D video datasets with 2D gaze target labels from gaze following tasks, our approach achieves the following key contributions: (i) Significant state-of-the-art improvements in within-domain and cross-domain generalization on unconstrained benchmarks like Gaze360 and GFIE, with notable cross-modal gains in video gaze estimation; (ii) Superior cross-domain performance on datasets such as MPIIFaceGaze and Gaze360 compared to frontal face methods. Code and pre-trained models will be released to the community.

3D视线估计弱监督多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。