arXiv:2509.17321cs.ROcs.CL2025-09被引 6

开源视觉语言模型在机器人任务进度预测上表现不足,仅达闭源模型70%水平

OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation

  • 基于视觉语言模型构建任务进度预测基准OpenGVL
  • 开源模型在复杂操作任务中仅达到闭源模型70%性能
  • 可自动筛选高质量机器人数据,适合数据集构建者使用

数据稀缺仍是制约机器人发展的主要瓶颈。尽管野外可用的机器人数据正呈指数增长,但可靠的任务完成时间预测能力有助于规模化自动标注与数据清洗。本文提出OpenGVL,一个针对包含机器人和人类执行者的多样化复杂操作任务的时序进展预测基准。评估了多个公开可用的开源基础模型,发现其在时序进展预测任务上的表现显著低于闭源模型,仅达到后者的约70%。此外,我们展示了OpenGVL在大规模机器人数据集质量评估中的实际应用价值,支持高效自动化数据筛选。相关基准与完整代码已开源至OpenGVL。

原文摘要 · Abstract (English)

Data scarcity remains one of the most limiting factors in driving progress in robotics. However, the amount of available robotics data in the wild is growing exponentially, creating new opportunities for large-scale data utilization. Reliable temporal task completion prediction could help automatically annotate and curate this data at scale. The Generative Value Learning (GVL) approach was recently proposed, leveraging the knowledge embedded in vision-language models (VLMs) to predict task progress from visual observations. Building upon GVL, we propose OpenGVL, a comprehensive benchmark for estimating task progress across diverse challenging manipulation tasks involving both robotic and human embodiments. We evaluate the capabilities of publicly available open-source foundation models, showing that open-source model families significantly underperform closed-source counterparts, achieving only approximately $70\%$ of their performance on temporal progress prediction tasks. Furthermore, we demonstrate how OpenGVL can serve as a practical tool for automated data curation and filtering, enabling efficient quality assessment of large-scale robotics datasets. We release the benchmark along with the complete codebase at \href{github.com/budzianowski/opengvl}{OpenGVL}.

机器人数据清洗视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。