arXiv:2508.13546cs.CV2025-08被引 1

无需眼动硬件,用软件预测用户视线,提升VR渲染效率

GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering

  • 用视觉变换器+LSTM融合场景与视线动态,纯软件预测视线
  • 误差仅3.83度,比传统方法快24%,且在各类场景中稳定表现
  • 适合希望低成本适配VR的开发者和应用厂商

视点渲染能显著降低虚拟现实应用的计算负担,通过将渲染质量集中在用户注视区域。当前方法依赖昂贵的基于硬件的眼动追踪系统,因成本高、校准复杂及硬件兼容性问题限制了广泛应用。本文提出GazeProphet,一种无需专用眼动追踪硬件的纯软件化视线预测方法。该方法结合球面视觉变压器处理360度VR场景,利用LSTM时序编码器捕捉视线序列模式,并通过多模态融合网络整合空间场景特征与时间视线动态,预测未来视线位置并输出置信度估计。在综合性VR数据集上的实验表明,GazeProphet实现3.83度的中位角误差,较传统显著性基线提升24%,且置信度校准可靠。该方法在不同空间区域和场景类型中保持一致性能,可在无额外硬件要求下实际部署于VR系统。统计分析证实各项指标改进均具显著性。结果表明,纯软件视线预测可用于VR视点渲染,使性能提升更易被各类VR平台与应用采用。

原文摘要 · Abstract (English)

Foveated rendering significantly reduces computational demands in virtual reality applications by concentrating rendering quality where users focus their gaze. Current approaches require expensive hardware-based eye tracking systems, limiting widespread adoption due to cost, calibration complexity, and hardware compatibility constraints. This paper presents GazeProphet, a software-only approach for predicting gaze locations in VR environments without requiring dedicated eye tracking hardware. The approach combines a Spherical Vision Transformer for processing 360-degree VR scenes with an LSTM-based temporal encoder that captures gaze sequence patterns. A multi-modal fusion network integrates spatial scene features with temporal gaze dynamics to predict future gaze locations with associated confidence estimates. Experimental evaluation on a comprehensive VR dataset demonstrates that GazeProphet achieves a median angular error of 3.83 degrees, outperforming traditional saliency-based baselines by 24% while providing reliable confidence calibration. The approach maintains consistent performance across different spatial regions and scene types, enabling practical deployment in VR systems without additional hardware requirements. Statistical analysis confirms the significance of improvements across all evaluation metrics. These results show that software-only gaze prediction can work for VR foveated rendering, making this performance boost more accessible to different VR platforms and apps.

VR渲染视线预测软件方案眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。