FovealNet提升眼动追踪精度,让虚拟现实渲染更高效清晰。
FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Optimized Foveated Rendering System Performance in Virtual Reality
- 采用事件触发裁剪与动态令牌修剪,减少64.8%无关像素计算
- 相比之前方法提速1.42倍,眼动追踪精度提升使渲染质量提高13%
- 支持多分辨率训练,适配不同实时渲染需求,降低系统开销
借助实时眼动追踪,缩放渲染技术可优化硬件效率并提升虚拟现实(VR)画质。该方法通过识别用户注视位置,仅在中央视觉区(即高视觉敏锐度的视锥区域)进行高分辨率渲染,而周边区域则以低分辨率处理。然而,现有基于深度学习的眼动追踪方案常存在长尾误差分布,导致追踪偏差,降低缩放渲染效果。本文提出FovealNet,一种先进AI驱动的眼动追踪框架,通过策略性增强追踪精度以优化系统性能。为降低算法实现成本,FovealNet引入基于事件的图像裁剪方法,消除超过64.8%的无关像素输入;同时采用简单有效的动态令牌修剪策略,在不损失精度的前提下即时移除冗余信息。此外,提出一种面向系统性能的多分辨率训练策略,使眼动追踪深度神经网络能适应不同运行时渲染配置,更有效地优化整体系统表现。评估结果表明,FovealNet相较以往方法至少实现1.42倍加速,并使缩放输出感知质量提升13%。
原文摘要 · Abstract (English)
Leveraging real-time eye-tracking, foveated rendering optimizes hardware efficiency and enhances visual quality virtual reality (VR). This approach leverages eye-tracking techniques to determine where the user is looking, allowing the system to render high-resolution graphics only in the foveal region-the small area of the retina where visual acuity is highest, while the peripheral view is rendered at lower resolution. However, modern deep learning-based gaze-tracking solutions often exhibit a long-tail distribution of tracking errors, which can degrade user experience and reduce the benefits of foveated rendering by causing misalignment and decreased visual quality. This paper introduces \textit{FovealNet}, an advanced AI-driven gaze tracking framework designed to optimize system performance by strategically enhancing gaze tracking accuracy. To further reduce the implementation cost of the gaze tracking algorithm, FovealNet employs an event-based cropping method that eliminates over $64.8\%$ of irrelevant pixels from the input image. Additionally, it incorporates a simple yet effective token-pruning strategy that dynamically removes tokens on the fly without compromising tracking accuracy. Finally, to support different runtime rendering configurations, we propose a system performance-aware multi-resolution training strategy, allowing the gaze tracking DNN to adapt and optimize overall system performance more effectively. Evaluation results demonstrate that FovealNet achieves at least $1.42\times$ speed up compared to previous methods and 13\% increase in perceptual quality for foveated output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。