用文字指令生成符合物体语义的双手抓握动作,解决数据少难题。
DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
- 基于对称性扩展单手数据,构建大规模双手抓握数据集
- 在少量标注物体上训练,却能泛化到未见物体
- 支持文本指令生成多样且语义一致的双手抓握方案
学习生成符合物体语义的双手抓握动作对实现稳定人机交互至关重要,但因数据稀缺而研究不足。现有抓握数据集多聚焦单手交互,且仅包含有限的语义部件标注。为此,我们提出一个名为 SymOpt 的流程,利用现有单手数据集并借助物体与手部对称性,构建大规模双手抓握数据集。在此基础上,我们设计了文本引导的双手抓握生成器 DHAGrasp,可为未见物体合成符合语义的双手抓握动作。该方法引入新颖的双手语义表示,采用两阶段架构,使模型能在少量标注物体上有效学习,并扩展至大量未分割数据。大量实验表明,本方法生成的抓握动作多样且语义一致,在抓握质量与未见物体泛化能力上均优于强基线。项目主页:https://quanzhou-li.github.io/DHAGrasp/
原文摘要 · Abstract (English)
Learning to generate dual-hand grasps that respect object semantics is essential for robust hand-object interaction but remains largely underexplored due to dataset scarcity. Existing grasp datasets predominantly focus on single-hand interactions and contain only limited semantic part annotations. To address these challenges, we introduce a pipeline, SymOpt, that constructs a large-scale dual-hand grasp dataset by leveraging existing single-hand datasets and exploiting object and hand symmetries. Building on this, we propose a text-guided dual-hand grasp generator, DHAGrasp, that synthesizes Dual-Hand Affordance-aware Grasps for unseen objects. Our approach incorporates a novel dual-hand affordance representation and follows a two-stage design, which enables effective learning from a small set of segmented training objects while scaling to a much larger pool of unsegmented data. Extensive experiments demonstrate that our method produces diverse and semantically consistent grasps, outperforming strong baselines in both grasp quality and generalization to unseen objects. The project page is at https://quanzhou-li.github.io/DHAGrasp/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。