arXiv:2609.05324cs.ROcs.AI2026-09

构建复杂空间与长程任务的机器人评估基准,检验视觉语言动作模型真实推理能力。

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

  • 设计双维度评测框架:精细空间推理与长程流程规划
  • 包含280种任务变体,覆盖10类场景与5个难度等级
  • 适合研究具身智能、机器人泛化与长程规划的学者

视觉-语言-动作(VLA)模型在语言控制的机器人操作中展现出显著进展。然而,现有数据集和评估标准多限定于预设场景下的任务完成,难以揭示模型在日益复杂的空间与流程环境中的推理能力。为此,我们提出 extbf{RoboSPA}( extbf{Robo}t extbf{S}patial- extbf{P}rocedural extbf{A}ssessment),一个大规模机器人操作数据集与评估基准,用于诊断VLA模型的具身推理能力。 exttt{RoboSPA}聚焦两个核心维度:细粒度空间推理与长时程流程规划,涵盖10类任务与56个基础任务。每项任务设置五个难度层级,共生成280种变体,逐步提升空间模糊性与流程复杂度。我们收集了跨多种机器人形态与多样场景的52.7万条轨迹。除二元成功率外, exttt{RoboSPA}引入诊断性指标以实现更细致评估。对代表性VLA模型的实验表明,当前系统在复杂空间关系理解、精准低层执行及高记忆负载规划方面仍表现不足。该结果确立了 exttt{RoboSPA}作为挑战性诊断基准的地位,推动更具能力、可靠且通用的具身智能体发展。数据与代码已公开于https://github.com/fanzhenxuan/RoboSPA。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

机器人具身智能评估基准长程规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。