用头动和视觉线索预测虚拟现实中的视线,无需眼动追踪
Gaze Prediction in Virtual Reality Without Eye Tracking Using Visual and Head Motion Cues
- 结合头显运动与视频显著性特征预测视线
- 在EHTask数据集上优于中心点、平均视线等基线方法
- 适合无眼动追踪硬件的VR应用,提升交互自然度
视线预测在虚拟现实(VR)应用中至关重要,可降低传感器延迟并支持如注视点渲染等计算密集型技术,这些技术依赖于对用户注意力的提前预测。然而,由于硬件限制或隐私顾虑,直接眼动追踪常不可用。为此,本文提出一种新型视线预测框架,结合头戴式显示器(HMD)运动信号与来自视频帧的视觉显著性线索。方法采用轻量级显著性编码器UniSal提取视觉特征,再与HMD运动数据融合,并通过时间序列预测模块进行未来视线方向预测。我们在EHTask数据集上评估了两种轻量级架构——TSMixer与LSTM,并在商用VR硬件上部署验证。结果表明,该方法持续优于中心点法(Center-of-HMD)与平均视线法(Mean Gaze),证明了在缺乏直接眼动追踪条件下,预测性视线建模能有效减少感知延迟,提升VR环境中的自然交互体验。
原文摘要 · Abstract (English)
Gaze prediction plays a critical role in Virtual Reality (VR) applications by reducing sensor-induced latency and enabling computationally demanding techniques such as foveated rendering, which rely on anticipating user attention. However, direct eye tracking is often unavailable due to hardware limitations or privacy concerns. To address this, we present a novel gaze prediction framework that combines Head-Mounted Display (HMD) motion signals with visual saliency cues derived from video frames. Our method employs UniSal, a lightweight saliency encoder, to extract visual features, which are then fused with HMD motion data and processed through a time-series prediction module. We evaluate two lightweight architectures, TSMixer and LSTM, for forecasting future gaze directions. Experiments on the EHTask dataset, along with deployment on commercial VR hardware, show that our approach consistently outperforms baselines such as Center-of-HMD and Mean Gaze. These results demonstrate the effectiveness of predictive gaze modeling in reducing perceptual lag and enhancing natural interaction in VR environments where direct eye tracking is constrained.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。