让机器人通过理解语言和物理接触生成更自然的抓取动作
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
- 用接触点作为中间表示,分步生成手部抓取姿态
- 在DexGYS数据集上成功率达67.14%,比顶尖方法高3.83个百分点
- 支持部分接触提示,实现对抓取过程的精细控制
语言驱动的灵巧抓取需理解任务语义、三维几何与复杂手物交互。现有方法直接从观测映射抓取参数,缺乏对物理交互的中间推理。我们提出DextER,一种基于接触的具身推理框架,用于多指操作。核心思想是预测手指链接在物体表面的接触位置,作为具身感知的中间表征,连接任务语义与物理约束。DextER自回归生成具身接触标记(指定哪根手指在何处接触物体表面),再生成抓取标记以编码手部姿态。在DexGYS数据集上,其抓取成功率达67.14%,优于当前最优方法3.83个百分点,意图对齐度提升96.4%。还可通过部分接触指定实现可控生成,实现对抓取合成的细粒度调控。
原文摘要 · Abstract (English)
Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have been applied to this problem, existing approaches directly map observations to grasp parameters without intermediate reasoning about physical interactions. We present DextER, Dexterous Grasp Generation with Embodied Reasoning, which introduces contact-based embodied reasoning for multi-finger manipulation. Our key insight is that predicting which hand links contact where on the object surface provides an embodiment-aware intermediate representation, bridging task semantics with physical constraints. DextER autoregressively generates embodied contact tokens specifying which finger links contact where on the object surface, followed by grasp tokens encoding the hand configuration. On DexGYS, DextER achieves 67.14% success rate, outperforming state-of-the-art by 3.83 p.p. with 96.4% improvement in intention alignment. We also demonstrate steerable generation through partial contact specification, providing fine-grained control over grasp synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。