端到端重建可动物体,效率提升30倍且精度更高
Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

- 通过隐空间几何学习,统一处理形状、部件与运动预测
- 在PartNet-Mobility上将误差降低20%,推理时间从9分钟缩至0.3秒
- 适合需要快速高精度可动物体重建的研究与工业应用
从稀疏图像中重建可动物体需恢复完整几何结构、可动部件及运动参数。现有方法通常分阶段进行几何重建、部件推理与运动估计,导致形状、活动部件与运动间一致性弱,且推理开销大。本文提出Artic-O,一种基于隐空间几何学习的端到端可动物体重建框架。不直接在图像或视图空间拟合几何,而是将多状态观测映射至预训练的隐几何空间,由冻结的流匹配解码器提供完整形状先验以恢复可见与被遮挡结构。为关联几何与运动,模型在图像引导的部件推理模块中融合视觉特征、几何隐变量与点级解码特征,实现活动部件分割与运动预测。通过几何-运动渐进式训练与解耦双阶段策略,平衡重建与部件级监督。在PartNet-Mobility数据集上,Artic-O实现强重建质量,相较基线方法LARM显著提升效率:Chamfer Distance降低20%,F-score提高,多数关节指标达到相当或更优水平,推理时间从9分钟降至约0.3秒/对象。
原文摘要 · Abstract (English)
Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry reconstruction, part reasoning, and articulation estimation into different stages. This separation can weaken consistency between shape, active parts, and motion, while also incurring substantial inference cost. We introduce Artic-O, an end-to-end, feed-forward framework for articulated object reconstruction via latent geometry learning. Instead of fitting geometry in image or view space, Artic-O maps sparse multi-state observations into a pretrained latent geometry space, where a frozen flow-matching decoder provides a complete-shape prior for recovering visible and occluded structures. To connect geometry with articulation, Artic-O fuses visual tokens, geometry latents, and point-wise decoder features in an image-grounded part-reasoning module for active-part segmentation and articulation prediction. We further train the model with a geometry-to-articulation curriculum and a decoupled two-pass strategy to balance reconstruction and part-level supervision. On PartNet-Mobility, Artic-O achieves strong reconstruction quality while being substantially more efficient than LARM, a strong prior method. It reduces Chamfer Distance, improves F-score, and achieves comparable or better articulation accuracy across most joint metrics, while reducing inference time from 9 minutes to about 0.3 seconds per object.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。