用扩散模型同步生成人体与物体互动的自然抓握动作。
Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model
- 统一用扩散模型建模人体、双手与物体运动的协同关系。
- 引入接触感知损失,使抓握姿态更符合物体空间位置。
- 适合动画、VR/AR和机器人领域需要真实人体交互的场景。
在动画、VR/AR和机器人等领域,生成高质量全身人机交互运动序列日益重要。该任务的核心挑战在于,面对不同尺寸和运动轨迹的复杂物体,需合理确定双手参与程度,同时保证抓握的真实感及全身动作协调性。现有方法或忽略手部精细抓握姿态,或仅建模静态抓握。本文提出一种简单而有效的框架,通过单一扩散模型联合建模身体、双手与给定物体运动序列之间的关系。为引导网络感知物体空间位置并学习更自然的抓握姿态,我们设计了新颖的接触感知损失,并引入数据驱动的精细引导机制。实验表明,本方法优于当前最先进方法,能生成合理的全身运动序列。
原文摘要 · Abstract (English)
Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in different sizes and their different motion trajectories, while ensuring strong grasping realism and guaranteeing the coordination of movement in all body parts. Contrasting with existing work, which either generates human interaction motion sequences without detailed hand grasping poses or only models a static grasping pose, we propose a simple yet effective framework that jointly models the relationship between the body, hands, and the given object motion sequences within a single diffusion model. To guide our network in perceiving the object's spatial position and learning more natural grasping poses, we introduce novel contact-aware losses and incorporate a data-driven, carefully designed guidance. Experimental results demonstrate that our approach outperforms the state-of-the-art method and generates plausible whole-body motion sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。