arXiv:2506.05576cs.RO2025-06被引 1

新数据集+新方法,让机器人零样本识别物体并抓取

TD-TOG Dataset: Benchmarking Zero-Shot and One-Shot Task-Oriented Grasping for Object Generalization

  • 用真实场景构建数据集,含物体掩码、功能区域和抓取点标注
  • 零样本识物+单样本学功能,多物场景抓取准确率达68.9%
  • 首次对比零样本与单样本在任务抓取中的泛化能力

任务导向抓取(TOG)是机器人执行任务的关键前置步骤,需预测目标物体上能完成特定任务的抓取位置。现有文献表明,尽管需求大,但可用于训练和评估的TOG数据集仍严重不足,且多为合成数据或存在掩码标注瑕疵。此外,多数TOG方法需依赖功能掩码、抓取点和物体掩码,但现有数据集通常仅提供部分标注。为此,我们提出顶视任务导向抓取(TD-TOG)数据集,包含1,449个真实世界RGB-D场景,涵盖30类30个子类别物体,均有人工标注的物体掩码、功能区域及平面矩形抓取点。数据集还包含一个用于评估模型区分物体子类能力的新测试集。为支持无需重训练即可处理未见物体的TOG方案,我们提出一种新框架Binary-TOG:利用零样本进行物体识别,单样本学习实现功能识别。在多物体场景中,Binary-TOG平均任务导向抓取准确率达68.9%。本文还提供了关于零样本与单样本学习在TOG中泛化能力的对比分析,以指导未来研究。

原文摘要 · Abstract (English)

Task-oriented grasping (TOG) is an essential preliminary step for robotic task execution, which involves predicting grasps on regions of target objects that facilitate intended tasks. Existing literature reveals there is a limited availability of TOG datasets for training and benchmarking despite large demand, which are often synthetic or have artifacts in mask annotations that hinder model performance. Moreover, TOG solutions often require affordance masks, grasps, and object masks for training, however, existing datasets typically provide only a subset of these annotations. To address these limitations, we introduce the Top-down Task-oriented Grasping (TD-TOG) dataset, designed to train and evaluate TOG solutions. TD-TOG comprises 1,449 real-world RGB-D scenes including 30 object categories and 120 subcategories, with hand-annotated object masks, affordances, and planar rectangular grasps. It also features a test set for a novel challenge that assesses a TOG solution's ability to distinguish between object subcategories. To contribute to the demand for TOG solutions that can adapt and manipulate previously unseen objects without re-training, we propose a novel TOG framework, Binary-TOG. Binary-TOG uses zero-shot for object recognition, and one-shot learning for affordance recognition. Zero-shot learning enables Binary-TOG to identify objects in multi-object scenes through textual prompts, eliminating the need for visual references. In multi-object settings, Binary-TOG achieves an average task-oriented grasp accuracy of 68.9%. Lastly, this paper contributes a comparative analysis between one-shot and zero-shot learning for object generalization in TOG to be used in the development of future TOG solutions.

任务抓取零样本数据集机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。