让虚拟物体自动摆出符合场景和物理的自然姿势
Towards Affordance-Aware Articulation Synthesis for Rigged Objects
- 用扩散模型生成符合上下文的交互信息
- 通过可微渲染实现不同骨骼结构的快速对齐
- 无需人工干预,几分钟内完成任意物体的姿势生成
拟人化物体在艺术创作中广泛应用,但使其灵活适应特定场景和姿态仍高度依赖经验艺术家的手动调整。本文提出A3Syn,解决开放域拟人物体的逼真、情境感知姿态合成问题。给定环境网格与文本提示,A3Syn可为互联网获取的任意骨骼结构物体生成合理关节参数。由于缺乏训练数据且不假设拓扑结构,研究采用2D图像修复扩散模型生成情境感知的交互信息,并结合可微渲染与语义对应关系,实现高效的骨骼对齐。该方法收敛稳定,生成耗时仅数分钟,在多种真实物体与场景组合上均能生成合理姿态。
原文摘要 · Abstract (English)
Rigged objects are commonly used in artist pipelines, as they can flexibly adapt to different scenes and postures. However, articulating the rigs into realistic affordance-aware postures (e.g., following the context, respecting the physics and the personalities of the object) remains time-consuming and heavily relies on human labor from experienced artists. In this paper, we tackle the novel problem and design A3Syn. With a given context, such as the environment mesh and a text prompt of the desired posture, A3Syn synthesizes articulation parameters for arbitrary and open-domain rigged objects obtained from the Internet. The task is incredibly challenging due to the lack of training data, and we do not make any topological assumptions about the open-domain rigs. We propose using 2D inpainting diffusion model and several control techniques to synthesize in-context affordance information. Then, we develop an efficient bone correspondence alignment using a combination of differentiable rendering and semantic correspondence. A3Syn has stable convergence, completes in minutes, and synthesizes plausible affordance on different combinations of in-the-wild object rigs and scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。