用连续对应嵌入生成逼真手物操作动作,实现物理合理且精细的交互。
ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion
- 基于3D物体轨迹,通过顶点级对应嵌入建模手物接触关系。
- 在扩散过程中引入残差修正,动态优化手部姿态生成结果。
- 支持单/双手抓取及刚性/可动物体,适合机器人操控与动画生成。
动态灵巧的物体操作极具挑战性,需同步手部运动与物体轨迹以实现流畅自然的物理合理交互。本文提出ManiDext,一种基于分层扩散框架的统一方法,根据3D物体轨迹生成手部操作与抓握姿态。核心洞察在于:准确建模交互中手物间的接触对应关系至关重要。为此,我们提出连续对应嵌入表示,在顶点层面精确刻画物体与手部之间的对应关系。该嵌入在手部网格上以自监督方式优化,其距离反映测地距离。框架首先在物体表面生成接触图与对应嵌入;随后,在第二阶段的手部姿态生成中,将迭代优化过程融入扩散过程:在每一步去噪中,引入当前手部姿态残差作为修正目标,引导网络纠正不准确姿态。该设计使生成与精修融合为统一框架,实验表明本方法能生成多种任务下(包括单/双手抓握、刚性与可动物体)高度真实且物理合理的运动序列。代码将用于研究目的。
原文摘要 · Abstract (English)
Dynamic and dexterous manipulation of objects presents a complex challenge, requiring the synchronization of hand motions with the trajectories of objects to achieve seamless and physically plausible interactions. In this work, we introduce ManiDext, a unified hierarchical diffusion-based framework for generating hand manipulation and grasp poses based on 3D object trajectories. Our key insight is that accurately modeling the contact correspondences between objects and hands during interactions is crucial. Therefore, we propose a continuous correspondence embedding representation that specifies detailed hand correspondences at the vertex level between the object and the hand. This embedding is optimized directly on the hand mesh in a self-supervised manner, with the distance between embeddings reflecting the geodesic distance. Our framework first generates contact maps and correspondence embeddings on the object's surface. Based on these fine-grained correspondences, we introduce a novel approach that integrates the iterative refinement process into the diffusion process during the second stage of hand pose generation. At each step of the denoising process, we incorporate the current hand pose residual as a refinement target into the network, guiding the network to correct inaccurate hand poses. Introducing residuals into each denoising step inherently aligns with traditional optimization process, effectively merging generation and refinement into a single unified framework. Extensive experiments demonstrate that our approach can generate physically plausible and highly realistic motions for various tasks, including single and bimanual hand grasping as well as manipulating both rigid and articulated objects. Code will be available for research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。