用视觉语言模型自动判断机器人任务执行是否正确且高质量
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots

- 基于视觉语言模型自动评估机器人任务正确性与质量
- 在1000+视频上测试,Gemini召回率更高,GPT精确率更高
- 能量化自身不确定性,适合自动化测试场景
机器人测试需判断其任务执行是否正确、可靠且高质量,这被称为测试圣杯问题。传统方法依赖特定任务的符号化断言和人工评估,耗时、主观且易出错。为此,我们提出VISOR,一种基于视觉语言模型(VLM)的自动化测试断言评估方法,可替代昂贵的人工评估。VISOR能自动判断任务正确性与质量,突破了现有符号化断言仅提供通过/失败判断、无法量化质量的局限。考虑到VLM固有的不确定性,VISOR还显式量化自身评估中的不确定性。我们在超过1000个视频上对两个VLM(GPT和Gemini)进行了四类机器人任务的评估。结果表明,Gemini具有更高召回率,而GPT具有更高精确率。但两者均显示不确定性与正确性相关性低,无法用不确定性预测任务正确性。
原文摘要 · Abstract (English)
Testing robots requires assessing whether they perform their intended tasks correctly, dependably, and with high quality, a challenge known as the test oracle problem in software testing. Traditionally, this assessment relies on task-specific symbolic oracles for task correctness and on human manual evaluation of robot behavior, which is time-consuming, subjective, and error-prone. To address this, we propose VISOR, a Vision-Language Model (VLM)-based approach for automated test oracle assessment that eliminates the need of expensive human evaluations. VISOR performs automated evaluation of task correctness and quality, addressing the limitations of existing symbolic test oracles, which are task-specific and provide pass/fail judgments without explicitly quantifying task quality. Given the inherent uncertainty in VLMs, VISOR also explicitly quantifies its own uncertainty during test assessments. We evaluated VISOR using two VLMs, i.e., GPT and Gemini, across four robotic tasks on over 1,000 videos. Results show that Gemini achieves higher recall while GPT achieves higher precision. However, both models show low correlation between uncertainty and correctness, which prevents using uncertainty as a correctness predictor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。