arXiv:2601.12948cs.CV2026-01中稿 · 3DV 2026

用扩散模型从单张图估计3D视线与姿态,提升精度。

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation

  • 基于扩散模型生成多组合理3D视线与姿态假设。
  • 在三个基准数据集上达到当前最优,超越依赖时序信息的方法。
  • 将视线视为距眼睛固定距离的额外关节,实现联合建模。

我们提出GazeD,一种新的3D gaze估计方法,可从单张RGB图像同时推断3D gaze和人体姿态。利用扩散模型处理不确定性的能力,GazeD基于输入图像提取的2D上下文信息,生成多个合理的3D gaze与姿态假设。具体而言,其去噪过程以2D姿态、主体周围环境及场景上下文为条件。GazeD还引入一种新视角:将3D gaze表示为距离双眼固定距离的附加身体关节,因其与姿态密切相关,故可在扩散过程中实现联合去噪。在三个基准数据集上的评估表明,GazeD在3D gaze估计任务中表现卓越,甚至优于依赖时序信息的方法。项目详情见 https://aimagelab.ing.unimore.it/go/gazed。

原文摘要 · Abstract (English)

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed.

3D视线估计扩散模型姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。