从单张RGBD图重建多类可动物体的零件级形状与运动结构
Detection Based Part-level Articulated Object Reconstruction from Single RGBD Image

- 通过检测零件再组合,实现跨类别零件级重建
- 在合成与真实数据上均优于已有方法,支持多样结构物体
- 适合需要精细运动建模的机器人、3D交互场景
我们提出一种端到端可训练的跨类别方法,从单张RGBD图像中重建多个人造可动物体,重点在于零件级形状重建、姿态与运动学估计。不同于依赖实例级隐空间学习的方法,本工作聚焦于具有预定义零件数量的人造可动物体。我们提出一种新型零件级表示方式,将实例表示为检测到的零件组合。尽管检测-分组策略能有效处理不同零件结构和数量的实例,但面临误检、零件大小尺度差异以及因端到端训练导致模型膨胀等问题。为此,我们提出:1)测试时基于运动学的零件融合,以提升检测性能并抑制误检;2)各向异性尺度归一化,以适应不同大小和尺度的零件形状学习;3)特征空间与输出空间间的平衡协同优化策略,在保持模型规模的同时提升零件检测能力。在合成与真实数据上的评估表明,该方法成功重建了以往方法无法处理的多样化结构多实例,并在形状重建和运动学估计上优于现有方法。
原文摘要 · Abstract (English)
We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from previous works that rely on learning instance-level latent space, focusing on man-made articulated objects with predefined part counts. Instead, we propose a novel alternative approach that employs part-level representation, representing instances as combinations of detected parts. While our detect-then-group approach effectively handles instances with diverse part structures and various part counts, it faces issues of false positives, varying part sizes and scales, and an increasing model size due to end-to-end training. To address these challenges, we propose 1) test-time kinematics-aware part fusion to improve detection performance while suppressing false positives, 2) anisotropic scale normalization for part shape learning to accommodate various part sizes and scales, and 3) a balancing strategy for cross-refinement between feature space and output space to improve part detection while maintaining model size. Evaluation on both synthetic and real data demonstrates that our method successfully reconstructs variously structured multiple instances that previous works cannot handle, and outperforms prior works in shape reconstruction and kinematics estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。