arXiv:2511.02468cs.HCcs.CV2025-11被引 1

利用头姿信息提升眼动数据缺失时的补全精度,实现更真实的视觉注意力重建。

HAGI++: Head-Assisted Gaze Imputation and Generation

  • 融合头姿与眼动的多模态扩散模型,学习两者间动态关联。
  • 在三个大数据集上优于传统插值与深度学习方法,速度分布更贴近真实人类行为。
  • 可扩展至手部动作,极端缺失下仍能生成合理眼动轨迹,适合真实场景应用。

移动端眼动追踪在现实世界与扩展现实(XR)环境中对捕捉人类视觉注意力至关重要,广泛应用于行为研究与人机交互。然而,因眨眼、瞳孔检测误差或光照变化导致的数据缺失,严重阻碍了后续分析。为此,我们提出HAGI++——首个利用头姿传感器整合眼动与头动相关性的多模态扩散模型。该方法采用基于Transformer的扩散模型,学习眼动与头姿表示间的跨模态依赖关系,并可轻松扩展以融入其他身体运动信号。在大规模Nymeria、Ego-Exo4D和HOT3D数据集上的评估表明,HAGI++在眼动补全任务中持续优于传统插值方法及基于深度学习的时间序列补全基线。统计分析显示,其生成的眼动速度分布与真实人类行为高度一致,确保了更高的真实性。此外,在100%眼动数据缺失的极端情况下,通过引入商用可穿戴设备获取的手腕运动信息,HAGI++超越了依赖全身动作捕捉的现有方法,实现纯眼动生成。本方法为真实场景下的完整、精确眼动记录开辟了新路径,具有显著提升眼动分析与交互性能的潜力。

原文摘要 · Abstract (English)

Mobile eye tracking plays a vital role in capturing human visual attention across both real-world and extended reality (XR) environments, making it an essential tool for applications ranging from behavioural research to human-computer interaction. However, missing values due to blinks, pupil detection errors, or illumination changes pose significant challenges for further gaze data analysis. To address this challenge, we introduce HAGI++ - a multi-modal diffusion-based approach for gaze data imputation that, for the first time, uses the integrated head orientation sensors to exploit the inherent correlation between head and eye movements. HAGI++ employs a transformer-based diffusion model to learn cross-modal dependencies between eye and head representations and can be readily extended to incorporate additional body movements. Extensive evaluations on the large-scale Nymeria, Ego-Exo4D, and HOT3D datasets demonstrate that HAGI++ consistently outperforms conventional interpolation methods and deep learning-based time-series imputation baselines in gaze imputation. Furthermore, statistical analyses confirm that HAGI++ produces gaze velocity distributions that closely match actual human gaze behaviour, ensuring more realistic gaze imputations. Moreover, by incorporating wrist motion captured from commercial wearable devices, HAGI++ surpasses prior methods that rely on full-body motion capture in the extreme case of 100% missing gaze data (pure gaze generation). Our method paves the way for more complete and accurate eye gaze recordings in real-world settings and has significant potential for enhancing gaze-based analysis and interaction across various application domains.

眼动补全多模态融合扩散模型真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。