用视觉语言模型反馈优化机器人抓取,提升成功率。
GraspCorrect: Robotic Grasp Correction via Vision-Language Model-Guided Feedback
- 通过视觉问答迭代生成抓取目标,结合任务约束与物理可行性筛选
- 在RLBench和CALVIN数据集上显著提升现有策略模型的抓取成功率
- 无需修改原模型,可即插即用,适合工业场景落地
尽管机器人操作取得显著进展,但实现稳定可靠的抓取仍是核心挑战,常成为复杂任务执行的瓶颈。我们分析发现,即使最先进的策略模型也频繁出现抓取不稳,导致失败。为此,提出GraspCorrect——一种通过视觉语言模型引导反馈来增强抓取性能的即插即用模块。该模块采用迭代式视觉问答框架,包含抓取引导提示(融入任务特定约束)与物体感知采样(确保物理可行的抓取候选)。通过不断生成中间视觉目标并转化为关节级动作,显著提升抓取稳定性,在RLBench和CALVIN数据集上持续改善现有策略模型的任务成功率达30%以上。
原文摘要 · Abstract (English)
Despite significant advancements in robotic manipulation, achieving consistent and stable grasping remains a fundamental challenge, often limiting the successful execution of complex tasks. Our analysis reveals that even state-of-the-art policy models frequently exhibit unstable grasping behaviors, leading to failure cases that create bottlenecks in real-world robotic applications. To address these challenges, we introduce GraspCorrect, a plug-and-play module designed to enhance grasp performance through vision-language model-guided feedback. GraspCorrect employs an iterative visual question-answering framework with two key components: grasp-guided prompting, which incorporates task-specific constraints, and object-aware sampling, which ensures the selection of physically feasible grasp candidates. By iteratively generating intermediate visual goals and translating them into joint-level actions, GraspCorrect significantly improves grasp stability and consistently enhances task success rates across existing policy models in the RLBench and CALVIN datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。