利用眼手头协同运动提升虚拟现实中的注视估计精度
HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination
- 通过眼手头协同性筛选优质训练样本,实现数据去噪
- 在HOT3D和ADT数据集上分别降低15.6%和6.0%的平均角度误差
- 适合需要高精度注视估计的元宇宙与人机交互应用
我们提出HOIGaze——一种基于学习的扩展现实(XR)中手-物体交互(HOI)场景下的注视估计新方法。该方法基于关键洞察:在手-物体交互过程中,眼睛、手和头部动作高度协调,这种协调性可用于识别对注视估计器训练最有价值的样本,从而有效去噪训练数据。与以往将所有样本视为等价的方法不同,本工作提出三点创新:1)一种分层框架,先识别当前视觉关注的手,再据此估计注视方向;2)一种融合头部与手-物体特征的跨模态Transformer注视估计器,特征由卷积神经网络和时空图卷积网络提取;3)一种新的眼-头协同损失,提升属于协调眼-头运动的样本权重。在HOT3D与Aria数字孪生(ADT)数据集上评估显示,其平均角度误差相比现有最优方法分别降低15.6%和6.0%。进一步实验表明,该方法在ADT上的基于注视的行为识别任务中也显著提升性能。结果表明眼-手-头协同蕴含丰富信息,为基于学习的注视估计开辟了新方向。
原文摘要 · Abstract (English)
We present HOIGaze - a novel learning-based approach for gaze estimation during hand-object interactions (HOI) in extended reality (XR). HOIGaze addresses the challenging HOI setting by building on one key insight: The eye, hand, and head movements are closely coordinated during HOIs and this coordination can be exploited to identify samples that are most useful for gaze estimator training - as such, effectively denoising the training data. This denoising approach is in stark contrast to previous gaze estimation methods that treated all training samples as equal. Specifically, we propose: 1) a novel hierarchical framework that first recognises the hand currently visually attended to and then estimates gaze direction based on the attended hand; 2) a new gaze estimator that uses cross-modal Transformers to fuse head and hand-object features extracted using a convolutional neural network and a spatio-temporal graph convolutional network; and 3) a novel eye-head coordination loss that upgrades training samples belonging to the coordinated eye-head movements. We evaluate HOIGaze on the HOT3D and Aria digital twin (ADT) datasets and show that it significantly outperforms state-of-the-art methods, achieving an average improvement of 15.6% on HOT3D and 6.0% on ADT in mean angular error. To demonstrate the potential of our method, we further report significant performance improvements for the sample downstream task of eye-based activity recognition on ADT. Taken together, our results underline the significant information content available in eye-hand-head coordination and, as such, open up an exciting new direction for learning-based gaze estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。