通过联合学习动态与结构,从单图生成可动3D物体
PWM-ArtGen: Part World Model for Articulated Object Generation

- 构建统一的部件世界模型,联合建模视觉动态与运动结构
- 在19.7k无标注图像对上训练,零样本泛化能力显著提升
- 适合做复杂可动物体生成的研究者和工业设计场景
从单张图像生成可动3D物体的关键挑战在于准确预测其运动结构。现有方法或直接从静态图像推断运动参数(缺乏动态部件关系),或通过单图生成视觉动态再估计参数,存在两步误差累积问题。此外,现有标注数据集规模小、多样性不足,限制了对复杂真实物体的泛化能力。为此,我们提出学习视觉动态与运动参数的联合分布,将可动物体视为动态系统,构建统一的部件世界模型PWM-ArtGen。为利用未标注数据,该模型耦合动作扩散与图像扩散,采用独立扩散步数,实现视觉分支的协同训练。我们进一步构建了一个包含19.7k张部件级图像对的逼真数据集,无运动标注,支持协同训练。实验表明,PWM-ArtGen在静止状态生成上显著优于基线,并展现出强大的零样本泛化能力,适用于分布外物体。
原文摘要 · Abstract (English)
The key challenge in articulated 3D object generation from a single image is accurately predicting the underlying kinematic structure. Existing methods either infer kinematic parameters directly from a static image that lacks dynamic part-level kinematic relationships, or estimate parameters from visual dynamics generated from a single image, which is prone to accumulated errors of two steps. Moreover, the limited scale and diversity of existing annotated datasets further hinder generalization to complex, real-world objects. To overcome these limitations, we propose to learn the joint distribution of visual dynamics and kinematic parameters. Recognizing that articulated objects can be formulated as dynamic systems, we propose a unified Part World Model called PWM-ArtGen. To leverage unannotated data, this model couples action diffusion and image diffusion with independent diffusion timesteps, which enables visual branch co-training. We further curate a photorealistic dataset of 19.7k part-level image pairs without kinematic annotations, to support co-training. Experiments demonstrate that PWM-ArtGen substantially outperforms existing baselines in the resting state and exhibits strong zero-shot generalization to out-of-distribution objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。