arXiv:2511.08007cs.CV2025-11AAAI被引 2

EAGLE统一2D与3D视觉定位,提升第一人称视角下目标检索精度。

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

  • 用外观与几何双记忆机制,融合元学习与几何追踪
  • 在Ego4D-VQ上达当前最优,显著提升长短期外观变化建模能力
  • 适合需要精准空间定位的智能穿戴与虚拟现实场景

第一人称视觉中的视觉查询定位对具身AI和VR/AR至关重要,但受相机运动、视角变化和外观差异影响仍具挑战。本文提出EAGLE框架,利用分段引导的外观感知元学习记忆(AMM)与几何感知定位记忆(GLM)协同工作,实现统一的2D-3D视觉查询定位。该记忆整合机制通过结构化外观与几何记忆库,存储高置信度检索样本,有效支持目标外观变化的长短时建模,实现精确轮廓分割与鲁棒空间区分,显著提升检索准确率。进一步结合VQL-2D输出与视觉几何驱动的Transformer(VGGT),高效实现2D与3D任务统一,支持快速精准的3D空间反投影。本方法在Ego4D-VQ基准上达到当前最优性能。

原文摘要 · Abstract (English)

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query localization in egocentric vision. Inspired by avian memory consolidation, EAGLE synergistically integrates segmentation guided by an appearance-aware meta-learning memory (AMM), with tracking driven by a geometry-aware localization memory (GLM). This memory consolidation mechanism, through structured appearance and geometry memory banks, stores high-confidence retrieval samples, effectively supporting both long- and short-term modeling of target appearance variations. This enables precise contour delineation with robust spatial discrimination, leading to significantly improved retrieval accuracy. Furthermore, by integrating the VQL-2D output with a visual geometry grounded Transformer (VGGT), we achieve a efficient unification of 2D and 3D tasks, enabling rapid and accurate back-projection into 3D space. Our method achieves state-ofthe-art performance on the Ego4D-VQ benchmark.

视觉定位第一人称视觉多模态记忆2D-3D统一

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。