单目相机下实时追踪2D/3D点,无需多视角或离线处理。
DynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction
- 用动态3D高斯分布在线重建场景并追踪点运动
- 在单帧图像上实现相机位姿与物体运动联合估计
- 无对应关系监督,首次实现单目在线点追踪基准
场景重建与运动追踪本质上相辅相成。追踪点可支持几何重建[14],而动态场景的几何重建又能实现点在时间上的3D追踪[24,39]。近期方法利用此机制将2D追踪提升至3D,以解决遮挡模糊问题[38]。然而,现有方法通常依赖离线处理或多视角相机,难以应用于机器人导航或混合现实等真实场景。本文提出动态单目在线重建(DynOMo),首次实现从非标定单目摄像头输入中在线完成2D与3D点追踪。基于3D高斯点云实现动态场景在线重建,扩展3D高斯以捕捉新内容和物体运动,并从单帧RGB图像中估计相机运动。该方法通过鲁棒图像特征重建与新颖的相似性增强正则化项,实现点轨迹自发生成,无需任何对应级监督。其性能达到现有方法水平,建立了首个单目未标定相机下的在线点追踪基准,旨在推动该领域发展,拓展实际应用范围。
原文摘要 · Abstract (English)
Reconstructing scenes and tracking motion are two sides of the same coin. Tracking points allow for geometric reconstruction [14], while geometric reconstruction of (dynamic) scenes allows for 3D tracking of points over time [24, 39]. The latter was recently also exploited for 2D point tracking to overcome occlusion ambiguities by lifting tracking directly into 3D [38]. However, above approaches either require offline processing or multi-view camera setups both unrealistic for real-world applications like robot navigation or mixed reality. We target the challenge of online 2D and 3D point tracking from unposed monocular camera input introducing Dynamic Online Monocular Reconstruction (DynOMo). We leverage 3D Gaussian splatting to reconstruct dynamic scenes in an online fashion. Our approach extends 3D Gaussians to capture new content and object motions while estimating camera movements from a single RGB frame. DynOMo stands out by enabling emergence of point trajectories through robust image feature reconstruction and a novel similarity-enhanced regularization term, without requiring any correspondence-level supervision. It sets the first baseline for online point tracking with monocular unposed cameras, achieving performance on par with existing methods. We aim to inspire the community to advance online point tracking and reconstruction, expanding the applicability to diverse real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。