arXiv:2604.04138cs.ROcs.AI2026-04中稿 · ed

用稀疏分类指导学习灵巧抓取,提升新物体泛化与可控性

Learning Dexterous Grasping from Sparse Taxonomy Guidance

  • 先基于场景和任务预测抓取分类,再生成连续指动
  • 在新物体上达到87.9%成功率,优于基线方法
  • 通过选择不同分类实现抓取策略调整,适合真实场景

灵巧操作需为物体和任务规划合适的抓取姿态,并通过多指协调控制执行。但为每个物体和任务指定密集的位姿或接触目标不切实际。而仅依赖任务奖励的端到端强化学习缺乏可控性,故障时难以干预。为此,我们提出GRIT,一种两阶段框架,从稀疏分类指导中学习灵巧控制。GRIT首先根据场景和任务上下文预测基于分类的抓取规范。在此稀疏指令下,策略生成连续指部运动,完成任务并保持预期抓取结构。结果显示,某些抓取分类对特定物体几何更有效。利用这一关系,GRIT在新物体上提升泛化能力,整体成功率达87.9%。此外,真实世界实验表明其具备可控性,可通过基于物体几何和任务意图的高层分类选择调整抓取策略。

原文摘要 · Abstract (English)

Dexterous manipulation requires planning a grasp configuration suited to the object and task, which is then executed through coordinated multi-finger control. However, specifying grasp plans with dense pose or contact targets for every object and task is impractical. Meanwhile, end-to-end reinforcement learning from task rewards alone lacks controllability, making it difficult for users to intervene when failures occur. To this end, we present GRIT, a two-stage framework that learns dexterous control from sparse taxonomy guidance. GRIT first predicts a taxonomy-based grasp specification from the scene and task context. Conditioned on this sparse command, a policy generates continuous finger motions that accomplish the task while preserving the intended grasp structure. Our result shows that certain grasp taxonomies are more effective for specific object geometries. By leveraging this relationship, GRIT improves generalization to novel objects over baselines and achieves an overall success rate of 87.9%. Moreover, real-world experiments demonstrate controllability, enabling grasp strategies to be adjusted through high-level taxonomy selection based on object geometry and task intent.

灵巧操作强化学习抓取策略可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。