通过融合多源数据提升眼动估计在跨域场景下的鲁棒性。
Hybrid-Domain Adaptative Representation Learning for Gaze Estimation
- 利用高质量近眼图像无监督对齐,分离出与眼神相关表征。
- 在三个数据集上分别达到5.02°、3.36°和9.26°的误差,性能领先。
- 适合关注跨域泛化能力的眼动估计研究者使用。
基于外观的眼动估计近年来取得显著进展,旨在从单张人脸图像中预测精确的3D注视方向。然而,多数方法在跨域评估中表现严重下降,主要受表情、佩戴物和图像质量等无关因素干扰。为此,本文提出一种新型混合域自适应表示学习框架(HARL),利用多源混合数据集学习鲁棒的眼动表征。具体而言,通过无监督域自适应方式,将低质量人脸图像与高质量近眼图像提取的特征对齐,从而解耦出与眼神相关的表征,几乎不增加计算或推理开销。此外,分析了头部姿态的影响,并设计了一种简单高效的稀疏图融合模块,挖掘注视方向与头部姿态间的几何约束,生成更密集且鲁棒的表征。在EyeDiap、MPIIFaceGaze和Gaze360数据集上的大量实验表明,该方法分别实现5.02°、3.36°和9.26°的误差,达到当前最优水平,并在跨数据集评估中表现优异。代码已公开于https://github.com/da60266/HARL。
原文摘要 · Abstract (English)
Appearance-based gaze estimation, aiming to predict accurate 3D gaze direction from a single facial image, has made promising progress in recent years. However, most methods suffer significant performance degradation in cross-domain evaluation due to interference from gaze-irrelevant factors, such as expressions, wearables, and image quality. To alleviate this problem, we present a novel Hybrid-domain Adaptative Representation Learning (shorted by HARL) framework that exploits multi-source hybrid datasets to learn robust gaze representation. More specifically, we propose to disentangle gaze-relevant representation from low-quality facial images by aligning features extracted from high-quality near-eye images in an unsupervised domain-adaptation manner, which hardly requires any computational or inference costs. Additionally, we analyze the effect of head-pose and design a simple yet efficient sparse graph fusion module to explore the geometric constraint between gaze direction and head-pose, leading to a dense and robust gaze representation. Extensive experiments on EyeDiap, MPIIFaceGaze, and Gaze360 datasets demonstrate that our approach achieves state-of-the-art accuracy of $\textbf{5.02}^{\circ}$ and $\textbf{3.36}^{\circ}$, and $\textbf{9.26}^{\circ}$ respectively, and present competitive performances through cross-dataset evaluation. The code is available at https://github.com/da60266/HARL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。