arXiv:2510.17801cs.ROcs.CV2025-10被引 23

评测大模型作为机器人认知核心的综合能力,覆盖理解、规划、故障分析等维度。

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

  • 构建多模态大模型作为机器人‘大脑’的评估框架,融合真实机器人数据。
  • 涵盖14项能力、25个任务,含6092个问答对,支持长周期推理评估。
  • 适合研究具身智能、机器人认知与大模型融合的学者使用。

构建能在动态、非结构化环境中感知、推理并行动的机器人仍是核心挑战。当前具身系统多采用双系统范式:系统2负责高层推理,系统1处理底层控制。我们将系统2称为具身大脑,即操作决策的认知核心。尽管评估该大脑至关重要,现有基准主要衡量执行成功率,且仅覆盖有限的高层认知与任务真实性。为此,我们提出RoboBench,一个用于评估多模态大语言模型(MLLMs)作为具身大脑的综合性基准。它涵盖五大维度:指令理解、感知推理、泛化规划、可及性预测和故障分析,共14项能力、25个任务和6,092个问答对。为提升真实性,数据源自大规模真实机器人数据及自建数据集,覆盖多样机器人形态、属性丰富的物体、多视角场景和记忆驱动导航。在规划方面,引入MLLM作为世界模拟器框架,评估预测计划是否能在物理与视觉约束下实现关键对象状态变化,从而更真实地评估长时程推理能力。对18个先进MLLMs的实验揭示其在隐含指令理解、时空推理、跨场景规划、细粒度可及性理解及故障诊断方面存在持续局限。我们进一步分析了具身认知能力与下游机器人控制之间的关系。RoboBench为量化高层认知提供了全面框架,并指导下一代MLLM向更稳健的机器人智能演进。

原文摘要 · Abstract (English)

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reasoning and System 1 handles low-level control. We refer to System 2 as the embodied brain, the cognitive core for decision-making in manipulation. Although evaluating this embodied brain is crucial, existing benchmarks mainly measure execution success or cover only limited aspects of high-level cognition and task realism. We introduce RoboBench, a benchmark for evaluating multimodal large language models (MLLMs) as embodied brains. RoboBench covers five dimensions: Instruction Comprehension, Perception Reasoning, Generalized Planning, Affordance Prediction, and Failure Analysis. It spans 14 capabilities, 25 tasks, and 6,092 QA pairs. To improve realism, it draws from large-scale real robotic data and in-house collection across diverse embodiments, attribute-rich objects, multi-view scenes, and memory-driven navigation. For planning, RoboBench introduces an MLLM-as-world-simulator framework that assesses whether predicted plans can achieve critical object-state changes under physical and visual constraints, enabling more faithful evaluation of long-horizon reasoning than symbolic matching. Experiments on 18 state-of-the-art MLLMs reveal persistent limitations in implicit instruction understanding, spatiotemporal reasoning, cross-scenario planning, fine-grained affordance understanding, and failure diagnosis. We further analyze how embodied cognitive abilities relate to downstream robotic control. RoboBench offers a comprehensive scaffold for quantifying high-level cognition and guiding next-generation MLLMs toward more robust robotic intelligence.

具身智能大模型评测机器人认知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。