用视觉思维链提升机器人抓取的推理能力,支持复杂环境下的多物体抓取。
VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- 引入视觉思维链机制,动态聚焦图像关键区域并生成可解释推理过程。
- 在167K合成数据和400+真实图像上训练,抓取成功率显著提升。
- 适合需要强泛化能力的机器人抓取场景,尤其在杂乱环境中表现优异。
机器人抓取是操作任务中最基础的一环,语言驱动的抓取因其实际交互潜力成为研究热点。然而,现有方法或缺乏足够推理与泛化能力,或依赖复杂模块化流程。当前抓取基础模型往往过度关注对话和对象语义,导致性能下降且仅限于单物体抓取。为在杂乱环境中保持强推理与泛化能力,我们提出VCoT-Grasp,一种端到端的抓取基础模型,通过引入视觉链式思维(Visual Chain-of-Thought)增强视觉理解。该模型采用多轮处理范式,动态关注视觉输入并生成可解释的推理轨迹。训练方面,我们构建并优化了大规模数据集VCoT-GraspSet,包含167,000张合成图像及超过136万次抓取标注,以及400余张真实图像和超过1200次抓取标注,均附带中间边界框。在该数据集及真实机器人上的实验表明,本方法显著提升了抓取成功率,并有效泛化至未见物体、背景和干扰物。
原文摘要 · Abstract (English)
Robotic grasping is one of the most fundamental tasks in robotic manipulation, and grasp detection/generation has long been the subject of extensive research. Recently, language-driven grasp generation has emerged as a promising direction due to its practical interaction capabilities. However, most existing approaches either lack sufficient reasoning and generalization capabilities or depend on complex modular pipelines. Moreover, current grasp foundation models tend to overemphasize dialog and object semantics, resulting in inferior performance and restriction to single-object grasping. To maintain strong reasoning ability and generalization in cluttered environments, we propose VCoT-Grasp, an end-to-end grasp foundation model that incorporates visual chain-of-thought reasoning to enhance visual understanding for grasp generation. VCoT-Grasp adopts a multi-turn processing paradigm that dynamically focuses on visual inputs while providing interpretable reasoning traces. For training, we refine and introduce a large-scale dataset, VCoT-GraspSet, comprising 167K synthetic images with over 1.36M grasps, as well as 400+ real-world images with more than 1.2K grasps, annotated with intermediate bounding boxes. Extensive experiments on both VCoT-GraspSet and real robot demonstrate that our method significantly improves grasp success rates and generalizes effectively to unseen objects, backgrounds, and distractors. More details can be found at https://zhanghr2001.github.io/VCoT-Grasp.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。