arXiv:2509.19524cs.AIcs.RO2025-09中稿 · CoRL被引 11

用视觉语言模型评估机器人操作每一步的成功率,让进步看得见。

Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation

  • 用VLM自动分析视频判断每一步操作是否成功
  • 输出每一步的独立成功率,揭示策略优劣细节
  • 框架轻量开源,适合各实验室复用和改进

机器人学习研究通常只报告单一的二元成功率(SR),掩盖了策略在多步操作任务中何时成功或失败。我们主张应常规化地报告子目标级结果:对每条轨迹输出一个按子目标划分的成功率向量,使部分能力可见(如抓取与倒出)。我们提出StepEval的蓝图——一种基于视觉语言模型(VLM)的低成本插件式评估框架,利用VLM作为自动化裁判,从记录的图像或视频中判断子目标成败。本工作不提出新基准或API,而是为可扩展、社区驱动的开源项目提供设计原则。在StepEval中,政策评估的核心产出是每子目标的成功率向量;同时考虑延迟或成本估计等其他指标,用于框架优化诊断,帮助社区在有真实标签时提升评估效率与准确性。框架保持模型无关性,支持单/多视角输入,足够轻量以跨实验室采纳。其核心贡献是共同方向:一个最小、可扩展的种子,鼓励开源协作,推动‘评分步骤而非仅最终目标’成为标准且可复现的实践。

原文摘要 · Abstract (English)

Robot learning papers typically report a single binary success rate (SR), which obscures where a policy succeeds or fails along a multi-step manipulation task. We argue that subgoal-level reporting should become routine: for each trajectory, a vector of per-subgoal SRs that makes partial competence visible (e.g., grasp vs. pour). We propose a blueprint for StepEval, a cost-aware plug-in evaluation framework that utilizes vision-language models (VLMs) as automated judges of subgoal outcomes from recorded images or videos. Rather than proposing new benchmarks or APIs, our contribution is to outline design principles for a scalable, community-driven open-source project. In StepEval, the primary artifact for policy evaluation is the per-subgoal SR vector; however, other quantities (e.g., latency or cost estimates) are also considered for framework-optimization diagnostics to help the community tune evaluation efficiency and accuracy when ground-truth subgoal success labels are available. We discuss how such a framework can remain model-agnostic, support single- or multi-view inputs, and be lightweight enough to adopt across labs. The intended contribution is a shared direction: a minimal, extensible seed that invites open-source contributions, so that scoring the steps, not just the final goal, becomes a standard and reproducible practice.

机器人子目标评估VLM可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。