从单张随意视频中恢复物体关节运动,无需复杂设备或标注。
Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video

- 用几何体拟合替代点追踪,避免跟踪失败
- 联合优化部件分割与关节参数,还原复杂运动结构
- 适合真实场景下有遮挡和剧烈运动的物体分析
从单目视频中恢复可动物体的3D运动关系是计算机视觉的基础挑战。现有方法依赖复杂的视频采集条件或长期点追踪、宽基线匹配等线索,但在严重遮挡、快速相机运动或局部特征弱时表现脆弱。基于学习的方法也难以泛化到训练外类别。本文提出一种类别无关的优化框架,将可动物体理解为几何原型拟合问题:用几何体作为代理表示,避免不稳定的点轨迹;设计新机制将原型组织成受转动/移动关节约束的连贯部件。该方法联合优化部件分割与关节参数,仅需一张随意拍摄的视频即可恢复复杂运动结构。引入可见性感知处理机制,有效应对真实数据中的部分观测与遮挡问题。同时构建了AiP-synth与AiP-real两个新基准,包含显著相机运动与重度遮挡,显著优于现有方法。
原文摘要 · Abstract (English)
Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long-term point tracking or wide-baseline matching, but are frequently brittle under severe occlusions, rapid camera ego-motion, or weak local features. Learning-based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category-agnostic optimization framework that treats articulated object understanding as a primitive-fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility-aware procedure handles partial observations and occlusions inherent to real-world data. We also propose the AiP-synth and AiP-real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods. Project page: https://aartykov.github.io/Articulation-in-Prime/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。