用颈挂摄像头估计用户视线,提升日常场景下眼动追踪精度
Egocentric Gaze Estimation via Neck-Mounted Camera
- 通过颈挂视角采集数据,构建首个该任务的数据集
- 引入视线越界分类辅助任务,显著提升模型性能
- 适合关注可穿戴设备与自然交互的研究者
本文提出颈挂视角眼动估计新任务,旨在从颈挂式摄像头视角推断用户视线。以往研究多聚焦头戴摄像头,而其他视角仍待探索。为此,我们收集了首个相关数据集,包含8名参与者在日常活动中录制的约4小时视频。在该数据集上评估基于Transformer的注视估计模型GLC,提出两种改进:辅助的视线越界分类任务和多视角联合学习方法,后者利用几何感知辅助损失同时训练头视图与颈视图模型。实验表明,引入视线越界分类可提升性能,但联合学习未带来增益。进一步分析结果并讨论其对颈挂式眼动估计的意义。
原文摘要 · Abstract (English)
This paper introduces neck-mounted view gaze estimation, a new task that estimates user gaze from the neck-mounted camera perspective. Prior work on egocentric gaze estimation, which predicts device wearer's gaze location within the camera's field of view, mainly focuses on head-mounted cameras while alternative viewpoints remain underexplored. To bridge this gap, we collect the first dataset for this task, consisting of approximately 4 hours of video collected from 8 participants during everyday activities. We evaluate a transformer-based gaze estimation model, GLC, on the new dataset and propose two extensions: an auxiliary gaze out-of-bound classification task and a multi-view co-learning approach that jointly trains head-view and neck-view models using a geometry-aware auxiliary loss. Experimental results show that incorporating gaze out-of-bound classification improves performance over standard fine-tuning, while the co-learning approach does not yield gains. We further analyze these results and discuss implications for neck-mounted gaze estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。