用奖励模型筛选伪标签,让眼动估计在真实场景中更通用。
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
- 用无标注数据+奖励模型评估伪标签可靠性
- 跨域测试准确率领先,零样本泛化到4个新数据集
- 适合做真实世界眼动估计的系统开发者
当前3D眼动估计方法在不同数据域间泛化能力差,主要因标注数据稀缺且多样性不足。本文提出OmniGaze,一种半监督3D眼动估计框架,利用来自多样真实环境的大规模无标注面部图像缓解领域偏移,提升野外泛化能力。首先构建包含不同面部特征、背景、光照、头部姿态和眼遮挡的无标注图像集;为拓展分布覆盖,采用标准伪标签策略,并设计奖励模型评估伪标签可靠性。该模型不仅使用3D方向向量,还结合预训练视觉编码器提取的视觉嵌入及多模态大语言模型生成的眼动视角语义线索,计算置信度分数;进而选择高质量伪标签并加权用于损失计算。大量实验表明,OmniGaze在五个数据集上均达到最优性能,涵盖同域与跨域设置;同时作为可扩展的数据引擎,在四个未见数据集上展现鲁棒零样本泛化能力。
原文摘要 · Abstract (English)
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to i) the scarcity of annotated datasets, and ii) the insufficient diversity of labeled data. In this work, we present OmniGaze, a semi-supervised framework for 3D gaze estimation, which utilizes large-scale unlabeled data collected from diverse and unconstrained real-world environments to mitigate domain bias and generalize gaze estimation in the wild. First, we build a diverse collection of unlabeled facial images, varying in facial appearances, background environments, illumination conditions, head poses, and eye occlusions. In order to leverage unlabeled data spanning a broader distribution, OmniGaze adopts a standard pseudo-labeling strategy and devises a reward model to assess the reliability of pseudo labels. Beyond pseudo labels as 3D direction vectors, the reward model also incorporates visual embeddings extracted by an off-the-shelf visual encoder and semantic cues from gaze perspective generated by prompting a Multimodal Large Language Model to compute confidence scores. Then, these scores are utilized to select high-quality pseudo labels and weight them for loss computation. Extensive experiments demonstrate that OmniGaze achieves state-of-the-art performance on five datasets under both in-domain and cross-domain settings. Furthermore, we also evaluate the efficacy of OmniGaze as a scalable data engine for gaze estimation, which exhibits robust zero-shot generalization on four unseen datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。