arXiv:2607.04610cs.RO2026-07中稿 · RSS 2026

构建机器人多场景评估基准,检验视觉语言模型真实应用能力

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

论文配图:RoboVista: Evaluating Vision Language Models for Diverse Robot Applications
图 1 · 摘自论文原文
  • 提出模块化RQA评估框架,解耦机器人决策组件
  • 涵盖39类任务的474个真实场景问答数据,覆盖农业到手术
  • 性能与实机表现高度相关,适合评估通用机器人推理能力

机器人在工业、农业等多样化应用场景中需应对不同机械结构、视觉条件和复杂规划。视觉语言模型(VLMs)为通用且可解释的机器人推理提供了潜力。但将VLMs适配多样机器人应用,需要对行为决策组件有模块化理解。传统基于遥控端到端数据集的评测难以捕捉这种结构。本文提出机器人问答(RQA)评估框架,并构建了基于真实机器人系统、研究论文及专家标注的基准RoboVista。该基准包含474个视觉问答实例,附有人工标注推理过程,覆盖农业、工业、家用、手术机器人、自动驾驶及开放数据集中的39种独特任务类型。在RoboVista上的实验表明,当前先进VLMs存在显著差距。物理机器人实验显示,其性能与实际任务执行效果强相关。

原文摘要 · Abstract (English)

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VLMs with diverse robot applications requires a modular understanding of the individual decision components that underlie robotic behavior. Capturing such structure is challenging for conventional robot benchmarks that are primarily based on teleoperated, end-to-end datasets. We propose Robot Question Answering (RQA), a modular evaluation framework and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations. RoboVista contains 474 Visual Question Answering (VQA) instances with human annotated reasoning and covers 39 unique task types in agricultural, industrial, domestic, surgical robotics, autonomous driving, and open robot datasets. Experiments on RoboVista show that state-of-the-art VLMs exhibit substantial gaps. Physical robot experiments suggest strong correlation between RoboVista performance and real-world task execution.

机器人视觉语言模型评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。