仅用4张图无姿态信息,就能重建可动物体的三维结构
PAOLI: Pose-free Articulated Object Learning from Sparse-view Images
- 先独立重建各关节,再通过变形场建立跨姿态对应关系
- 在仅四视图无姿态条件下实现高精度物体建模
- 适合缺乏相机位姿数据的实物扫描与逆向建模场景
我们提出一种新方法,仅需少量视角图像(每关节最少4张)且无需已知相机位姿,即可建模可动物体。现有方法依赖密集多视角观测和真实相机姿态。本方法首先利用最新的稀疏视角3D重建技术,独立重建每个关节;随后学习一个形变场,在不同姿态间建立稠密对应关系。采用渐进式解耦策略分离静态与运动部件,实现相机运动与物体运动的鲁棒分离。最后通过自监督损失联合优化几何、外观与运动学,强制保证跨视角和跨姿态的一致性。在标准基准和真实场景上的实验表明,该方法在远弱于现有方法输入条件的情况下,仍能生成精确且细节丰富的可动物体表示。
原文摘要 · Abstract (English)
We present a methodology to model articulated objects using a sparse set of images with unknown poses. Current methods require dense multi-view observations and ground-truth camera poses. Our approach operates with as few as four views per articulation and no camera supervision. Our central insight is to first solve a robust correspondence and alignment problem between unaligned reconstructions, before part motions can be analyzed. We first reconstruct each articulation independently using recent advances in sparse-view 3D reconstruction, then learn a deformation field that establishes dense correspondences across poses. A progressive disentanglement strategy further separates static from moving parts, enabling robust separation of camera and object motion. Finally, we optimize geometry, appearance, and kinematics jointly with a self-supervised loss that enforces cross-view and cross-pose consistency. Experiments on the standard benchmark and real-world examples demonstrate that our method produces accurate and detailed articulated object representations under significantly weaker input assumptions than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。