用推理+强化学习提升机器人抓取的场景理解能力
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
- 通过思维链引导实现先验推理与空间定位
- 在多个基准数据集上超越现有最优方法
- 适合复杂语言指令下的机器人抓取任务
我们提出AffordanceGrasp-R1,一种基于推理的物体抓取语义分割框架,结合思维链(CoT)冷启动策略与强化学习,增强推理能力与空间定位精度。同时,重构抓取流程:从全局场景点云生成抓取候选,并通过指令条件化的功能掩码进行筛选。大量实验表明,AffordanceGrasp-R1在多个基准数据集上持续优于当前最优(SOTA)方法;真实机器人抓取测试进一步验证其在复杂语言控制操作场景下的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
We introduce AffordanceGrasp-R1, a reasoning-driven affordance segmentation framework for robotic grasping that combines a chain-of-thought (CoT) cold-start strategy with reinforcement learning to enhance deduction and spatial grounding. In addition, we redesign the grasping pipeline to be more context-aware by generating grasp candidates from the global scene point cloud and subsequently filtering them using instruction-conditioned affordance masks. Extensive experiments demonstrate that AffordanceGrasp-R1 consistently outperforms state-of-the-art (SOTA) methods on benchmark datasets, and real-world robotic grasping evaluations further validate its robustness and generalization under complex language-conditioned manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。