用头动数据预测用户视线,提升移动端VR渲染效率
GazeProphetV2: Head-Movement-Based Gaze Prediction Enabling Efficient Foveated Rendering on Mobile VR
- 融合视线历史、头动和场景信息,动态加权关键信号
- 跨场景测试准确率达93.1%,未来1-3帧预测更准
- 无需昂贵眼动设备,适合移动端VR系统优化
在虚拟现实环境中预测用户视线行为仍是一大挑战,对渲染优化与界面设计具有重要意义。本文提出一种多模态方法,结合时间序列视线模式、头动数据与视觉场景信息,通过门控融合机制与跨模态注意力,自适应地根据上下文重要性权重组合三类信号。基于覆盖22个VR场景、包含530万条注视样本的数据集评估显示,融合多模态信息的预测精度优于单一数据流。结果表明,将历史视线轨迹、头部朝向与场景内容结合,可有效提升未来1至3帧的预测准确率。跨场景泛化测试中,验证准确率达93.1%,且预测视线轨迹具有良好的时间一致性。该研究深化了对虚拟环境中注意力机制的理解,为渲染优化、交互设计及用户体验评估提供了可行方案。该方法推动了无需高成本眼动硬件即可预判用户注意力的高效虚拟现实系统发展。
原文摘要 · Abstract (English)
Predicting gaze behavior in virtual reality environments remains a significant challenge with implications for rendering optimization and interface design. This paper introduces a multimodal approach to VR gaze prediction that combines temporal gaze patterns, head movement data, and visual scene information. By leveraging a gated fusion mechanism with cross-modal attention, the approach learns to adaptively weight gaze history, head movement, and scene content based on contextual relevance. Evaluations using a dataset spanning 22 VR scenes with 5.3M gaze samples demonstrate improvements in predictive accuracy when combining modalities compared to using individual data streams alone. The results indicate that integrating past gaze trajectories with head orientation and scene content enhances prediction accuracy across 1-3 future frames. Cross-scene generalization testing shows consistent performance with 93.1% validation accuracy and temporal consistency in predicted gaze trajectories. These findings contribute to understanding attention mechanisms in virtual environments while suggesting potential applications in rendering optimization, interaction design, and user experience evaluation. The approach represents a step toward more efficient virtual reality systems that can anticipate user attention patterns without requiring expensive eye tracking hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。