arXiv:2502.02307cs.CV2025-02被引 15

用大规模自监督预训练提升眼神估计跨域泛化能力

UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training

  • 基于野外人脸数据自监督预训练,构建通用眼神估计模型
  • 在跨数据集测试中显著提升泛化性能,减少对标注数据依赖
  • 适合需要低资源跨域眼神估计的视觉系统开发者

尽管数十年来在数据收集和模型架构方面持续研究,现有眼神估计模型在不同数据域间仍面临严重泛化挑战。近年来自监督预训练在多种视觉任务中表现出色,但在眼神估计领域尚未被充分探索。本文首次提出UniGaze,利用大规模野外人脸数据集进行自监督预训练,以实现通用眼神估计。通过系统性分析,我们明确了有效预训练的关键因素。实验表明,为语义任务设计的自监督方法在眼神估计中失效,而我们精心设计的预训练流程能持续提升跨域性能。在具有挑战性的跨数据集评估及新设定(如留一数据集、联合数据集)下,验证了UniGaze显著改善多域泛化能力,同时大幅降低对昂贵标注数据的依赖。源代码与模型已开源。

原文摘要 · Abstract (English)

Despite decades of research on data collection and model architectures, current gaze estimation models encounter significant challenges in generalizing across diverse data domains. Recent advances in self-supervised pre-training have shown remarkable performances in generalization across various vision tasks. However, their effectiveness in gaze estimation remains unexplored. We propose UniGaze, for the first time, leveraging large-scale in-the-wild facial datasets for gaze estimation through self-supervised pre-training. Through systematic investigation, we clarify critical factors that are essential for effective pretraining in gaze estimation. Our experiments reveal that self-supervised approaches designed for semantic tasks fail when applied to gaze estimation, while our carefully designed pre-training pipeline consistently improves cross-domain performance. Through comprehensive experiments of challenging cross-dataset evaluation and novel protocols including leave-one-dataset-out and joint-dataset settings, we demonstrate that UniGaze significantly improves generalization across multiple data domains while minimizing reliance on costly labeled data. source code and model are available at https://github.com/ut-vision/UniGaze.

眼神估计自监督学习跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。