提升复杂环境下的视线估计精度,通过超分辨率和双注意力机制融合头眼信息。
DHECA-SuperGaze: Dual Head-Eye Cross-Attention and Super-Resolution for Unconstrained Gaze Estimation
- 用双分支网络分别处理眼和高分辨率头部图像,结合交叉注意力增强特征
- 在Gaze360和GFIE数据集上,静态与动态场景下误差分别降低0.48°~3.00°
- 修复了Gaze360数据集中关键标注错误,显著提升模型泛化能力
无约束视线估计旨在识别个体在非受控环境中注视的方向。该技术广泛应用于驾驶分心监测、考试监考、软件无障碍功能等场景。然而,现有系统在真实环境下仍面临挑战,主要源于野外图像分辨率低,以及对头眼交互建模不足。本文提出DHECA-SuperGaze,一种基于深度学习的方法,通过超分辨率(SR)与双头眼交叉注意力(DHECA)模块提升视线预测性能。采用双分支卷积主干网络分别处理眼部与多尺度超分辨率头部图像,提出的DHECA模块通过交叉注意力实现头眼特征的双向优化。此外,发现并修正了当前最多样且广泛应用的视线数据集Gaze360中的关键标注错误。在Gaze360和GFIE数据集上的评估显示,本方法在静态配置下将角误差(AE)分别降低0.48°(Gaze360)和2.95°(GFIE),在时序设置下分别降低0.59°(Gaze360)和3.00°(GFIE)。跨数据集测试表明,静态与时序场景下误差均提升超过1.53°(Gaze360)和3.99°(GFIE),验证了方法的强泛化能力。
原文摘要 · Abstract (English)
Unconstrained gaze estimation is the process of determining where a subject is directing their visual attention in uncontrolled environments. Gaze estimation systems are important for a myriad of tasks such as driver distraction monitoring, exam proctoring, accessibility features in modern software, etc. However, these systems face challenges in real-world scenarios, partially due to the low resolution of in-the-wild images and partially due to insufficient modeling of head-eye interactions in current state-of-the-art (SOTA) methods. This paper introduces DHECA-SuperGaze, a deep learning-based method that advances gaze prediction through super-resolution (SR) and a dual head-eye cross-attention (DHECA) module. Our dual-branch convolutional backbone processes eye and multiscale SR head images, while the proposed DHECA module enables bidirectional feature refinement between the extracted visual features through cross-attention mechanisms. Furthermore, we identified critical annotation errors in one of the most diverse and widely used gaze estimation datasets, Gaze360, and rectified the mislabeled data. Performance evaluation on Gaze360 and GFIE datasets demonstrates superior within-dataset performance of the proposed method, reducing angular error (AE) by 0.48° (Gaze360) and 2.95° (GFIE) in static configurations, and 0.59° (Gaze360) and 3.00° (GFIE) in temporal settings compared to prior SOTA methods. Cross-dataset testing shows improvements in AE of more than 1.53° (Gaze360) and 3.99° (GFIE) in both static and temporal settings, validating the robust generalization properties of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。