用统一模型同时生成手与物体的自然抓握姿势。
Joint Diffusion for Universal Hand-Object Grasp Generation
- 单个扩散模型联合建模手与物体的潜在表示
- 在未见物体上仍能生成多样且合理的抓握姿势
- 适合动画与机器人抓取任务,尤其对新物体泛化好
预测并生成人手对物体的抓握姿势对动画与机器人任务至关重要。本文提出联合手-物扩散模型(JHOD),通过统一的潜在空间建模手与物体,利用手物抓握数据学习生成合理抓握。为增强对多样化物体形状的泛化能力,模型还引入大规模物体数据集学习包容性物体潜在嵌入。无论是否给定物体作为条件,该模型均可无条件或条件化生成抓握。相比仅依赖手物抓握数据的方法,本方法因使用更丰富的物体数据训练,在处理未见物体时表现更优。定量与定性实验表明,生成的手部抓握在视觉合理性与多样性上均表现良好,且能有效泛化到新物体形状。
原文摘要 · Abstract (English)
Predicting and generating human hand grasp over objects is critical for animation and robotic tasks. In this work, we focus on generating both the hand and objects in a grasp by a single diffusion model. Our proposed Joint Hand-Object Diffusion (JHOD) models the hand and object in a unified latent representation. It uses the hand-object grasping data to learn to accommodate hand and object to form plausible grasps. Also, to enforce the generalizability over diverse object shapes, it leverages large-scale object datasets to learn an inclusive object latent embedding. With or without a given object as an optional condition, the diffusion model can generate grasps unconditionally or conditional to the object. Compared to the usual practice of learning object-conditioned grasp generation from only hand-object grasp data, our method benefits from more diverse object data used for training to handle grasp generation more universally. According to both qualitative and quantitative experiments, both conditional and unconditional generation of hand grasp achieves good visual plausibility and diversity. With the extra inclusiveness of object representation learned from large-scale object datasets, the proposed method generalizes well to unseen object shapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。