arXiv:2605.07188cs.CV2026-05

PicoEyes统一建模眼动,支持无标定、重戴等复杂场景。

PicoEyes: Unified Gaze Estimation Framework for Mixed Reality with a Large-Scale Multi-View Dataset

论文配图:PicoEyes: Unified Gaze Estimation Framework for Mixed Reality with a Large-Scale Multi-View Dataset
图 1 · 摘自论文原文
  • 端到端联合预测3D眼参数、视线、深度图等多属性
  • 在多种设置下超越现有学术与工业方法性能
  • 适用于混合现实中的鲁棒眼动追踪,适合设备开发者

我们提出 PicoEyes,一种统一的眼动估计框架,可直接从单目或双目输入中预测所有关键眼动属性,包括3D眼参数、眼区分割、光学轴、视轴和深度图。该框架同时解决标定、眼动预测与设备姿态变化问题,并通过联合估计眼参数与深度图实现端到端的3D眼重建。此外,我们构建了一个大规模多视角近眼数据集,涵盖多样条件下(训练、测试、重戴测试、标定)的完整2D与3D标注。大量实验表明,PicoEyes在无标定、标定、重戴后标定及预测等多种场景下均显著优于当前学术与工业级眼动追踪方法。本工作为混合现实应用建立了实用、端到端的鲁棒且通用的眼动估计范式。

原文摘要 · Abstract (English)

We present PicoEyes, a unified gaze estimation framework that directly predicts all key attributes of gaze, including 3D eye parameters, eye-region segmentation, optical axis, visual axis, and depth maps, from either monocular or binocular inputs. The framework simultaneously addresses calibration, gaze forecasting, and varying device postures, while also supporting 3D eye reconstruction via joint estimation of eye parameters and depth maps in an end-to-end manner. In addition, we introduce a large-scale multi-view near-eye dataset containing comprehensive 2D and 3D annotations under diverse conditions, including train, test, rewear-test, and calibration sessions. Extensive experiments demonstrate that PicoEyes achieves state-ofthe-art performance, consistently outperforming both academic and industrial gaze tracking methods across nocalibration, calibration, rewear-after-calibration, and forecasting settings. This work establishes a practical, end-toend paradigm for robust and generalizable gaze estimation in mixed reality (MR) applications.

眼动估计混合现实端到端多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。