arXiv:2603.08021cs.ROcs.CV2026-03被引 1

用扩散模型生成符合物体特性与指令意图的自然抓握姿态

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

  • 基于扩散模型融合物体几何、空间可操作性与文本指令
  • 在多个数据集上抓握质量、语义准确率显著提升
  • 适合虚拟现实、机器人抓取等需要自然交互的场景

生成能准确反映物体几何形状和用户指定交互语义的人类抓握姿态,对于增强现实/虚拟现实及具身人工智能中的自然人机交互至关重要。现有语义抓握方法在3D物体表示与文本指令之间存在巨大模态鸿沟,且缺乏明确的空间或语义约束,导致生成结果常不物理可行或语义不一致。本文提出AffordGrasp,一种基于扩散的框架,可生成物理稳定且语义忠实的高精度人类抓握姿态。我们设计了一种可扩展的标注流程,自动为手-物交互数据集添加细粒度结构化语言标签以捕捉交互意图。在此基础上,AffordGrasp结合了感知可操作性的手部姿态隐空间表示与双条件扩散过程,实现对物体几何、空间可操作性及指令语义的联合推理。分布调整模块进一步确保物理接触一致性与语义对齐。我们在四个由HO-3D、OakInk、GRAB和AffordPose衍生的指令增强基准上评估该方法,结果表明其在抓握质量、语义准确性和多样性方面均显著优于现有最先进方法。

原文摘要 · Abstract (English)

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches struggle with the large modality gap between 3D object representations and textual instructions, and often lack explicit spatial or semantic constraints, leading to physically invalid or semantically inconsistent grasps. In this work, we present AffordGrasp, a diffusion-based framework that produces physically stable and semantically faithful human grasps with high precision. We first introduce a scalable annotation pipeline that automatically enriches hand-object interaction datasets with fine-grained structured language labels capturing interaction intent. Building upon these annotations, AffordGrasp integrates an affordance-aware latent representation of hand poses with a dual-conditioning diffusion process, enabling the model to jointly reason over object geometry, spatial affordances, and instruction semantics. A distribution adjustment module further enforces physical contact consistency and semantic alignment. We evaluate AffordGrasp across four instruction-augmented benchmarks derived from HO-3D, OakInk, GRAB, and AffordPose, and observe substantial improvements over state-of-the-art methods in grasp quality, semantic accuracy, and diversity.

抓握生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。