提出首个面向人机交互的零样本眼神估计基准,揭示现有方法在动态场景下的脆弱性。
Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction

- 构建包含50+受试者、3000+视频的大规模动态眼神数据集
- 发现所有方法在低头看时均失效,纯眼球模型表现最稳健
- 强调数据多样性比复杂模型更关键,适合机器人视觉研究者
零样本外观式3D眼神估计虽具成本优势,但在人机交互(HRI)场景中的可靠性尚不明确。现有基准常忽略动态相机视角与移动目标等核心HRI条件,且跨数据集评估存在显著复杂度差距。为此,我们推出Gaze4HRI,一个大规模数据集(50+受试者,3,000+视频,600,000+帧),用于评估眼神估计模型在光照变化、头部-视线冲突以及摄像头与注视目标运动等关键条件下的表现。实验表明,所有方法至少在一种条件下失效,低头注视是普遍失败点;仅ETH-X-Gaze训练的PureGaze在其余条件中保持稳定。结果挑战了当前对时空建模与Transformer架构的过度关注,表明数据多样性是零样本鲁棒性的首要驱动力,而如PureGaze的自对抗损失等净化机制可进一步提升性能。本研究为从业者提供实用指南,并重塑未来研究方向。数据与代码已公开。
原文摘要 · Abstract (English)
While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability in Human-Robot Interaction (HRI) settings remains uncertain. Existing benchmarks frequently overlook fundamental HRI conditions, such as dynamic camera viewpoints and moving targets in video. Furthermore, current cross-dataset evaluations often suffer from a complexity gap, where methods trained on diverse datasets are tested on significantly smaller and less varied sets, failing to assess true robustness. To bridge these gaps, we introduce Gaze4HRI, a large-scale dataset (50+ subjects, 3,000+ videos, 600,000+ frames) designed to evaluate state-of-the-art performance against critical HRI variables: illumination, head-gaze conflict, as well as the motion of camera and gaze target in video. Our benchmark reveals that all evaluated methods fail in at least one condition, identifying steeply-downward gaze as a universal failure point. Notably, PureGaze trained on the ETH-X-Gaze dataset uniquely maintains resilience across all other conditions. These results challenge the recent focus in the literature on complex spatial-temporal modeling and Transformer-based architectures. Instead, our findings suggest that extensive data diversity, as exemplified by the ETH-X-Gaze dataset, serves as the primary driver of zero-shot robustness in unconstrained environments, while resilience-enhancing frameworks, such as PureGaze's self-adversarial loss for gaze feature purification, provide a substantial further improvement. Ultimately, this study establishes a rigorous benchmark that provides practical guidelines for practitioners as well as reshaping future research. The dataset and codes are available at https://gazeforhri.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。