MAGE模型同时预测6自由度眼动方向,提升人机交互中的眼动分析精度。
MAGE: A Multi-task Architecture for Gaze Estimation with an Efficient Calibration Module
- 多任务架构融合方向与位置特征,通过专用信息流和多个解码器联合建模。
- 在MPIIFaceGaze、EYEDIAP和自建IMRGaze数据集上均达到当前最佳性能。
- 提出易用校准模块(Easy-Calibration),无需屏幕即可快速适配个体差异。
眼球注视可提供人类心理活动的丰富信息,在人机交互(HRI)领域备受关注。然而,现有注视估计方法仅能预测注视方向或屏幕上的注视点(PoG),难以实现三维空间中完整的六自由度(6-DoF)注视分析。此外,个体间眼形与结构差异也制约了方法的泛化能力。本文提出MAGE——一种带高效校准模块的多任务注视估计架构,用于实现实用化六自由度注视分析。基础模型从面部图像中编码方向与位置特征,通过专用信息流和多个解码器输出注视结果。为降低个体差异影响,提出新型校准模块Easy-Calibration,仅需用户特定数据即可高效微调模型,且无需屏幕。实验表明,该方法在公开数据集MPIIFaceGaze、EYEDIAP及自建数据集IMRGaze上均达到当前最优性能。
原文摘要 · Abstract (English)
Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interaction (HRI). However, existing gaze estimation methods merely predict either the gaze direction or the Point-of-Gaze (PoG) on the screen, failing to provide sufficient information for a comprehensive six Degree-of-Freedom (DoF) gaze analysis in 3D space. Moreover, the variations of eye shape and structure among individuals also impede the generalization capability of these methods. In this study, we propose MAGE, a Multi-task Architecture for Gaze Estimation with an efficient calibration module, to predict the 6-DoF gaze information that is applicable for the real-word HRI. Our basic model encodes both the directional and positional features from facial images, and predicts gaze results with dedicated information flow and multiple decoders. To reduce the impact of individual variations, we propose a novel calibration module, namely Easy-Calibration, to fine-tune the basic model with subject-specific data, which is efficient to implement without the need of a screen. Experimental results demonstrate that our method achieves state-of-the-art performance on the public MPIIFaceGaze, EYEDIAP, and our built IMRGaze datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。