用扩散模型重建可动物体,兼顾结构与运动关系。
KineDiff3D: Kinematic-Aware Diffusion for Category-Level Articulated Object Shape Reconstruction and Generation
- 通过关节感知的VAE将几何、关节角、部件分割编码到潜在空间
- 双扩散模型分别恢复姿态参数与生成完整形状潜码
- 迭代优化确保形变符合关节约束,适合复杂可动物体建模
可动物体(如笔记本、抽屉)因多部件结构和可变关节配置,在单视图3D重建与位姿估计中面临挑战。为此,我们提出KineDiff3D:一种统一框架,从单视图输入重建类别级可动物体并估计其位姿。首先,通过新型关节感知变分自编码器(KA-VAE)将完整几何(SDF)、关节角度和部件分割编码至结构化潜在空间;其次,采用两个条件扩散模型:一个用于回归全局姿态(SE(3))与关节参数,另一个用于从部分观测生成具备运动感知的潜在代码;最后,引入迭代优化模块,通过最小化Chamfer距离双向精修重建精度与关节参数,同时保持运动约束。在合成、半合成及真实数据集上的实验表明,该方法能准确重建可动物体并估计其运动属性。
原文摘要 · Abstract (English)
Articulated objects, such as laptops and drawers, exhibit significant challenges for 3D reconstruction and pose estimation due to their multi-part geometries and variable joint configurations, which introduce structural diversity across different states. To address these challenges, we propose KineDiff3D: Kinematic-Aware Diffusion for Category-Level Articulated Object Shape Reconstruction and Generation, a unified framework for reconstructing diverse articulated instances and pose estimation from single view input. Specifically, we first encode complete geometry (SDFs), joint angles, and part segmentation into a structured latent space via a novel Kinematic-Aware VAE (KA-VAE). In addition, we employ two conditional diffusion models: one for regressing global pose (SE(3)) and joint parameters, and another for generating the kinematic-aware latent code from partial observations. Finally, we produce an iterative optimization module that bidirectionally refines reconstruction accuracy and kinematic parameters via Chamfer-distance minimization while preserving articulation constraints. Experimental results on synthetic, semi-synthetic, and real-world datasets demonstrate the effectiveness of our approach in accurately reconstructing articulated objects and estimating their kinematic properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。