arXiv:2601.05243cs.ROcs.CV2026-01被引 2

仅用一次人类示范,就能让机器人学会抓握新物体。

Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration

  • 基于对应关系生成多样仿真数据,实现从单次示范迁移抓握策略。
  • 在多个类别物体上实现泛化抓握,性能显著优于现有方法。
  • 适合研究机器人灵巧操作与少样本学习的学者参考。

灵巧机器人手的功能性抓握是实现工具使用和复杂操作的关键能力,但长期以来受限于大规模数据集稀缺以及模型缺乏语义与几何推理的整合。本文提出 CorDex 框架,仅需一次人类示范即可从合成数据中鲁棒地学习新物体的灵巧功能抓握。核心是一个基于对应关系的数据引擎,根据人类示范生成同类别多样化物体实例,通过对应估计将专家抓握转移至生成物体,并通过优化进行适应。基于生成数据,我们设计了一个融合视觉与几何信息的多模态预测网络,采用局部-全局融合模块与重要性感知采样机制,实现高效鲁棒的灵巧抓握预测。在多种物体类别上的大量实验表明,CorDex 能有效泛化至未见物体实例,显著超越当前最优基线方法。

原文摘要 · Abstract (English)

Functional grasping with dexterous robotic hands is a key capability for enabling tool use and complex manipulation, yet progress has been constrained by two persistent bottlenecks: the scarcity of large-scale datasets and the absence of integrated semantic and geometric reasoning in learned models. In this work, we present CorDex, a framework that robustly learns dexterous functional grasps of novel objects from synthetic data generated from just a single human demonstration. At the core of our approach is a correspondence-based data engine that generates diverse, high-quality training data in simulation. Based on the human demonstration, our data engine generates diverse object instances of the same category, transfers the expert grasp to the generated objects through correspondence estimation, and adapts the grasp through optimization. Building on the generated data, we introduce a multimodal prediction network that integrates visual and geometric information. By devising a local-global fusion module and an importance-aware sampling mechanism, we enable robust and computationally efficient prediction of functional dexterous grasps. Through extensive experiments across various object categories, we demonstrate that CorDex generalizes well to unseen object instances and significantly outperforms state-of-the-art baselines.

灵巧抓握少样本学习仿真数据多模态网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。