仅用一张图生成可动物体,还原真实外观与运动结构。
SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects
- 用扩散模型学习单视角下部件形状与运动的合理变化
- 从结构到细节分步生成,提升真实感与输入一致性
- 适合3D资产快速建模,尤其家庭可动物品
我们解决从单张图像生成家用可动物体3D资产的挑战。现有方法或需多视角多状态输入,或仅能粗略控制生成过程,限制了建模的可扩展性与实用性。本文提出一种新方法,仅需任意视角下的静止物体图像,即可生成视觉上与输入一致的可动物体。为捕捉单视图带来的部件形状与运动模糊,设计了一个学习几何与运动合理变体的扩散模型。针对多领域属性结构化数据生成的复杂性,构建从高层结构到几何细节的粗到精生成流程,使用部件连接图与部件抽象作为代理。实验表明,该方法在生成物真实感、与输入图像的相似度及重建质量上显著优于现有最先进方法。
原文摘要 · Abstract (English)
We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation process. These limitations hinder the scalability and practicality for articulated object modeling. In this work, we propose a method to generate articulated objects from a single image. Observing the object in resting state from an arbitrary view, our method generates an articulated object that is visually consistent with the input image. To capture the ambiguity in part shape and motion posed by a single view of the object, we design a diffusion model that learns the plausible variations of objects in terms of geometry and kinematics. To tackle the complexity of generating structured data with attributes in multiple domains, we design a pipeline that produces articulated objects from high-level structure to geometric details in a coarse-to-fine manner, where we use a part connectivity graph and part abstraction as proxies. Our experiments show that our method outperforms the state-of-the-art in articulated object creation by a large margin in terms of the generated object realism, resemblance to the input image, and reconstruction quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。