新基准GCA-Bench测试机器人复杂抓取的多步推理能力。
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

- 构建包含场景推理与语义约束的复杂抓取任务评测集
- 现有方法在复杂场景下成功率低于70%
- 适合研究具身智能、通用机器人抓取的学者
稳健的机器人抓取仍是复杂现实应用中的根本挑战。尽管大规模模型在机器人任务推理方面展现出潜力,但现有抓取基准主要聚焦于孤立的视觉抓取姿态检测,未能涵盖需多步推理与语义理解的任务复杂性。为此,我们提出GCA-Bench,一个包含复杂动作抓取场景的基准,涉及场景级推理与语义约束。该基准支持在相同条件下评估近期大模型表现。我们实现了从传统抓取检测流程到端到端学习方法的多种基线。实证研究表明,在复杂抓取场景中成功率低于70%,凸显当前方法的关键局限。此外,我们提出了新的评估指标,分析关键失败模式,并为开发更鲁棒、泛化更强的抓取策略提供指导。
原文摘要 · Abstract (English)
Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning in robotic tasks. However, existing benchmarks for grasping primarily focus on isolated, visual-based grasp pose detection, failing to capture the complexity of grasping tasks that require multi-step reasoning and semantic understanding during execution. To address this gap, we propose GCA-Bench, a benchmark featuring challenging \textit{grasping with complex action} scenarios that involve both scene-level reasoning and semantic constraints. GCA-Bench enables the evaluation of recent large foundation models under the same settings. To demonstrate the effectiveness of our new benchmark, we implement a diverse set of baselines, ranging from traditional grasp detection pipelines to end-to-end learning methods. Empirical studies achieve success rates below 70\% on complex grasping scenarios, underscoring critical limitations. In addition, we propose new evaluation metrics, analyze critical failure models, and provide insights to guide the development of more robust and generalizable grasping strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。