arXiv:2410.08530cs.CVcs.MM2024-10中稿 · ACM Multimedia 202…被引 9

无需训练即可追踪第一人称视频中所有3D物体,精度显著提升。

Ego3DT: Tracking Every 3D Object in Ego-centric Videos

论文配图:Ego3DT: Tracking Every 3D Object in Ego-centric Videos
图 1 · 摘自论文原文
  • 基于相邻帧信息动态构建第一人称3D场景
  • 在新数据集上实现1.04至2.90倍的HOTA提升
  • 适合研究具身智能与第一人称视觉追踪的学者

具身智能的研究日益重视第一人称视角。然而,由于视角变化剧烈,准确定位和追踪第一人称视频中的物体仍是一大挑战。本文提出Ego3DT,一种零样本的3D重建与追踪框架,可对第一人称环境中的所有物体进行检测与分割。利用相邻帧信息,Ego3DT通过预训练的3D场景重建模型动态构建第一人称视角的3D场景。此外,我们设计了动态分层关联机制,以生成稳定可靠的3D追踪轨迹。在两个新构建的数据集上的大量实验表明,该方法在HOTA指标上达到1.04x–2.90x的提升,验证了其在多样化第一人称场景下的鲁棒性与准确性。

原文摘要 · Abstract (English)

The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily due to the substantial variability in viewing angles. Addressing this issue, this paper introduces a novel zero-shot approach for the 3D reconstruction and tracking of all objects from the ego-centric video. We present Ego3DT, a novel framework that initially identifies and extracts detection and segmentation information of objects within the ego environment. Utilizing information from adjacent video frames, Ego3DT dynamically constructs a 3D scene of the ego view using a pre-trained 3D scene reconstruction model. Additionally, we have innovated a dynamic hierarchical association mechanism for creating stable 3D tracking trajectories of objects in ego-centric videos. Moreover, the efficacy of our approach is corroborated by extensive experiments on two newly compiled datasets, with 1.04x - 2.90x in HOTA, showcasing the robustness and accuracy of our method in diverse ego-centric scenarios.

第一人称视觉3D追踪具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。