从第一视角视频中实时追踪3D物体,提升动态场景理解能力
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

- 通过2D分割掩码融合点运动评分与体素合并策略,实现3D轨迹关联
- 在ADT数据集上PCL指标提升11%,优于最强基线
- 支持稀疏深度输入和交互引导关联,适合真实复杂环境应用
从第一视角视频理解3D场景是机器人与自主导航的基础,但快速视角变化和部分遮挡使得构建结构化表征极具挑战。现有3D跟踪与场景图构建方法多针对显式交互或假设静态场景,难以捕捉复杂动态。我们提出EgoTrack3D,一个直接从第一视角RGB视频重建并维护动态3D场景表示的模块化框架。该框架将2D分割掩码提升至全局3D坐标系,结合基于点的运动评分机制与体素融合启发式策略实现物体轨迹关联。EgoTrack3D能长期保持高精度表示,在Aria Digital Twin(ADT)数据集上相较最强基线实现11%的正确位置比例(PCL)提升,覆盖静态与动态物体的持续3D跟踪。此外,为验证系统在退化条件下的鲁棒性,我们用稀疏3D边界框估计替代密集深度图,并引入交互引导的动态关联,使EgoTrack3D在观测噪声下仍能维持准确空间表征。
原文摘要 · Abstract (English)
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static scenes, limiting their ability to capture complex dynamics. We introduce EgoTrack3D, a modular framework that reconstructs and maintains a dynamic 3D scene representation directly from egocentric RGB video. The framework lifts 2D segmentation masks into a global 3D coordinate frame, using a point-based motion scoring mechanism alongside a voxel-based merging heuristic to associate object tracks. EgoTrack3D maintains accurate representations over time, achieving an 11% improvement in percentage of correct locations (PCL) relative to the strongest baseline on the Aria Digital Twin (ADT) dataset, while addressing the more general setting of persistent 3D tracking for both static and dynamic objects. Furthermore, to demonstrate the system's robustness under degraded conditions that simulate real-world deployment constraints, we replace dense depth maps with sparse 3D bounding box estimation and integrate interaction-guided dynamic association, enabling EgoTrack3D to maintain accurate spatial representations despite noisy observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。