提升无约束场景下眼神估计的泛化能力,适配戴眼镜、弱光等复杂条件。
Real-time Appearance-based Gaze Estimation for Open Domains
- 通过合成眼镜、口罩和多光照增强图像多样性,扩展数据分布。
- 采用多任务学习缓解不同数据集间标签偏差,尤其在俯仰角方向提升精度。
- 轻量模型仅用1%参数达SOTA性能,适合移动端实时运行。
基于外观的眼神估计(AGE)在受控环境下表现优异,但在实际开放场景中存在显著泛化差距,尤其在佩戴面部配件和弱光条件下表现不佳。我们归因于图像多样性不足与跨数据集标签一致性差,特别是俯仰角方向。为此,提出一种无需额外人工标注的鲁棒AGE框架:首先通过合成眼镜、口罩及多样化光照扩充图像流形;其次将眼神回归重构为多任务学习,引入多视角监督对比学习(SupCon)、离散标签分类与眼区分割作为辅助目标,以缓解异质标签偏差。为严格验证,构建新基准数据集,专门评估复杂条件下的鲁棒性。所提基于MobileNet的轻量模型性能媲美当前最优(UniGaze-H),参数量不足其1%,支持高保真、实时眼神追踪,适用于移动设备。
原文摘要 · Abstract (English)
Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly those involving facial wearables and poor lighting conditions. We attribute this failure to two core factors: limited image diversity and inconsistent label fidelity across different datasets, especially along the pitch axis. To address these, we propose a robust AGE framework that enhances generalization without requiring additional human-annotated data. First, we expand the image manifold via an ensemble of augmentation techniques, including synthesis of eyeglasses, masks, and varied lighting. Second, to mitigate the impact of anisotropic inter-dataset label deviation, we reformulate gaze regression as a multi-task learning problem, incorporating multi-view supervised contrastive (SupCon) learning, discretized label classification, and eye-region segmentation as auxiliary objectives. To rigorously validate our approach, we curate new benchmark datasets designed to evaluate gaze robustness under challenging conditions, a dimension largely overlooked by existing evaluation protocols. Our MobileNet-based lightweight model achieves generalization performance competitive with the state-of-the-art (SOTA) UniGaze-H, while utilizing less than 1\% of its parameters, enabling high-fidelity, real-time gaze tracking on mobile devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。