预测第一人称视角下人的3D视觉关注范围,提升AR/VR和辅助技术体验。
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
- 将2D视觉预测升级为3D场景中的视觉关注区域建模。
- 在364.6万样本数据集上实现领先性能,支持未来视觉范围精准预测。
- 适用于增强现实、智能助行系统等需理解人类行为意图的场景。
人类基于内在意图持续感知并互动于周围环境,而视觉感知在引导行为中起核心作用,但其在第一人称视角下的预测仍较少被研究,尤其对AR/VR与辅助技术意义重大。本文提出EgoSpanLift,将第一人称视觉关注范围从2D图像平面拓展至3D空间,通过转换SLAM生成的关键点为适配凝视的几何结构,并提取体积化视觉区域。结合3D U-Net与单向变压器,实现时空融合,高效预测3D网格中的未来视觉关注区域。同时,构建包含364.6万样本的多感官第一人称数据基准测试集。实验表明,该方法在2D凝视预测与3D定位任务上均优于现有基线,且无需额外2D训练即可达到相近2D表现。
原文摘要 · Abstract (English)
People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and contact-based interaction, forecasting human visual perception itself remains less explored despite its fundamental role in guiding human actions and its implications for AR/VR and assistive technologies. We address the challenge of egocentric 3D visual span forecasting, predicting where a person's visual perception will focus next within their three-dimensional environment. To this end, we propose EgoSpanLift, a novel method that transforms egocentric visual span forecasting from 2D image planes to 3D scenes. EgoSpanLift converts SLAM-derived keypoints into gaze-compatible geometry and extracts volumetric visual span regions. We further combine EgoSpanLift with 3D U-Net and unidirectional transformers, enabling spatio-temporal fusion to efficiently predict future visual span in the 3D grid. In addition, we curate a comprehensive benchmark from raw egocentric multisensory data, creating a testbed with 364.6K samples for 3D visual span forecasting. Our approach outperforms competitive baselines for egocentric 2D gaze anticipation and 3D localization while achieving comparable results even when projected back onto 2D image planes without additional 2D-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。