用3D眼球先验实现零样本眼动追踪,无需新数据采集
GazePrior: Zero-Shot AR/VR Eye Tracking via Learned 3D Gaze Reconstruction

- 基于人类眼球分布构建3D先验模型,支持跨设备数据重建
- 合成数据在真实度与精度上媲美实测数据,零样本性能领先
- 适合快速部署新头显设备的眼动系统,节省大量标注成本
眼动追踪是高级AR/VR应用的基础技术。但为每个新设备训练眼动模型面临挑战:真实数据收集成本高且耗时,现有合成数据生成方法又缺乏真实性。为此,我们提出一种数据驱动的3D眼球先验模型,刻画不同身份、注视方向和光照条件下人眼的分布特性。该模型名为GazePrior,可对以往设备采集的稀疏标注数据进行3D重建,并渲染至任意目标设备摄像头视角。本方法在不增加成本的前提下,生成具有真实感、多样性和真值精度的合成数据。实验表明,使用该合成数据训练的眼动模型,在零样本场景下优于此前方法,精度更高、鲁棒性更强。
原文摘要 · Abstract (English)
Eye tracking (ET) is a foundational technology for advanced AR/VR applications. However, training ET models for every new ET device is challenging: real data collection is costly and time-consuming, while existing synthetic data generation methods lack realism. To remove the need for additional data collection while maintaining data quality, we introduce a data-driven 3D prior that models the distribution of human eyes across diverse identities, gaze directions, and light settings. This model, which we coin GazePrior, then enables sparse-input 3D reconstruction of annotated data collected with previous ET devices, which can in turn be rendered from the cameras of any target ET device. Our approach synthesizes data with the realism, diversity and ground-truth accuracy of real data collection without its prohibitive costs. Our experiments demonstrate that ET models trained with our synthesized data outperform previous zero-shot methods, achieving higher accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。