arXiv:2505.09698cs.ROcs.AI2025-05被引 27

评测视觉语言模型在低级机器人操作中的推理能力,填补了该领域基准空白。

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

  • 构建多维度基准ManipBench,评估模型对物体交互与柔体操作的理解
  • 33个模型测试显示性能差异大,且与真实操作任务趋势高度相关
  • 当前模型表现远低于人类水平,适合研究低级机器人决策的学者参考

视觉语言模型(VLMs)凭借其常识推理能力革新了人工智能与机器人领域。在机器人操作中,VLMs 主要用于高层规划,但近期研究也开始关注其低级推理能力——即对精确机器人运动的决策能力。然而,当前社区缺乏一个清晰统一的基准来评估 VLMs 在低级机器人操作中的表现。为此,我们提出新基准 ManipBench,用于评估 VLMs 在多个维度上的低级操作推理能力,包括物体-物体交互和柔体对象操作。我们对 33 个代表性 VLMs(来自 10 个模型家族,含不同规模变体)进行了全面测试。结果表明,模型在任务间表现差异显著,且与真实世界操作任务的趋势存在强相关性。同时,模型表现仍与人类理解水平存在明显差距。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have revolutionized artificial intelligence and robotics due to their commonsense reasoning capabilities. In robotic manipulation, VLMs are used primarily as high-level planners, but recent work has also studied their lower-level reasoning ability, which refers to making decisions about precise robot movements. However, the community currently lacks a clear and common benchmark that can evaluate how well VLMs can aid low-level reasoning in robotics. Consequently, we propose a novel benchmark, ManipBench, to evaluate the low-level robot manipulation reasoning capabilities of VLMs across various dimensions, including how well they understand object-object interactions and deformable object manipulation. We extensively test 33 representative VLMs across 10 model families on our benchmark, including variants to test different model sizes. Our evaluation shows that the performance of VLMs significantly varies across tasks, and there is a strong correlation between this performance and trends in our real-world manipulation tasks. It also shows that there remains a significant gap between these models and human-level understanding. See our website at: https://manipbench.github.io.

机器人操作视觉语言模型基准评测低级推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。