从视频中自动推断物体关节结构,支持动态视角和部分遮挡。
Articulated Object Estimation in the Wild
- 结合点跟踪与因子图优化,从RGB-D视频中估计关节轨迹和轴线。
- 在新数据集Arti4D上优于传统与学习型基线方法。
- 适合做机器人操作、场景理解的研究者使用。
理解可动物体的3D运动对机器人场景理解、移动操作和运动规划至关重要。以往方法多依赖受控环境,假设固定相机视角或直接观测物体状态,难以适应真实复杂场景。人类却能通过观察他人操作轻松推断物体结构。受此启发,我们提出ArtiPoint框架,可在动态相机运动和部分可观测条件下,直接从原始RGB-D视频中估计可动部件轨迹与关节轴线。该方法融合深度点跟踪与因子图优化,实现鲁棒估计。为推动该领域研究,我们构建了首个以第一视角记录真实场景中可动物体交互的Arti4D数据集,包含关节标签与真实相机位姿。我们在Arti4D上对多种经典与基于学习的基线进行基准测试,验证了ArtiPoint的优越性能。代码与数据集已公开于https://artipoint.cs.uni-freiburg.de。
原文摘要 · Abstract (English)
Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming either fixed camera viewpoints or direct observations of various object states, which tend to fail in more realistic unconstrained environments. In contrast, humans effortlessly infer articulation by watching others manipulate objects. Inspired by this, we introduce ArtiPoint, a novel estimation framework that can infer articulated object models under dynamic camera motion and partial observability. By combining deep point tracking with a factor graph optimization framework, ArtiPoint robustly estimates articulated part trajectories and articulation axes directly from raw RGB-D videos. To foster future research in this domain, we introduce Arti4D, the first ego-centric in-the-wild dataset that captures articulated object interactions at a scene level, accompanied by articulation labels and ground-truth camera poses. We benchmark ArtiPoint against a range of classical and learning-based baselines, demonstrating its superior performance on Arti4D. We make code and Arti4D publicly available at https://artipoint.cs.uni-freiburg.de.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。