arXiv:2506.06677cs.ROcs.CV2025-06NeurIPS被引 32

构建大规模长时程机器人操作评估基准,推动视觉语言模型的深层推理能力

RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation

  • 用GPT生成任务指令并分解为复杂子任务序列,构建仿真数据集
  • 设计高低层协同框架,支持长周期规划与动态环境适应
  • 聚焦长期规划、反思与记忆能力,适合研究高级智能机器人系统者

视觉语言模型(VLM)的进步使指令驱动的机器人系统具备更强泛化能力。然而,现有工作多集中于反应式系统1策略,未能充分发挥VLM在语义推理与长时程规划中的优势。这类系统2能力——以目标为导向、深思熟虑的思维模式——因现有基准在时间跨度和结构复杂度上的局限而未被充分探索。为此,我们提出RoboCerebra,一个用于评估长时程机器人操作中高层推理能力的大规模基准。该基准包含:(1) 通过自上而下流程构建的大规模仿真数据集,涵盖家庭环境中更长的任务时序与多样化子任务序列;(2) 高层VLM规划器与低层视觉-语言-动作(VLA)控制器相结合的分层框架;(3) 通过结构化系统1-系统2交互,评估规划、反思与记忆能力的评测协议。任务指令由GPT生成并分解为子任务序列,人类操作员在仿真中执行,获得含动态物体变化的高质量轨迹。相比以往基准,RoboCerebra具有显著更长的动作序列和更密集的标注。我们进一步将先进VLM作为系统2模块进行基准测试,分析其在关键认知维度上的表现,推动更强大、更具泛化能力的机器人规划器发展。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs' strengths in semantic reasoning and long-horizon planning. These System 2 capabilities-characterized by deliberative, goal-directed thinking-remain under explored due to the limited temporal scale and structural complexity of current benchmarks. To address this gap, we introduce RoboCerebra, a benchmark for evaluating high-level reasoning in long-horizon robotic manipulation. RoboCerebra includes: (1) a large-scale simulation dataset with extended task horizons and diverse subtask sequences in household environments; (2) a hierarchical framework combining a high-level VLM planner with a low-level vision-language-action (VLA) controller; and (3) an evaluation protocol targeting planning, reflection, and memory through structured System 1-System 2 interaction. The dataset is constructed via a top-down pipeline, where GPT generates task instructions and decomposes them into subtask sequences. Human operators execute the subtasks in simulation, yielding high-quality trajectories with dynamic object variations. Compared to prior benchmarks, RoboCerebra features significantly longer action sequences and denser annotations. We further benchmark state-of-the-art VLMs as System 2 modules and analyze their performance across key cognitive dimensions, advancing the development of more capable and generalizable robotic planners.

机器人操作长时程规划视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。