评估视觉语言机器人执行质量与决策信心,超越传统成功与否的简单判断。
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- 引入13项专用于机器人的不确定性与质量指标
- 908次任务中多项指标与专家评价高度相关
- 可区分失败任务中的执行质量高低,适合无成功标签场景
视觉语言动作(VLA)机器人融合视觉感知、自然语言理解与动作规划,自主完成具身任务。当前评估多依赖任务成功率这一二元标准,难以反映执行质量与模型决策信心。本文针对三款先进VLA模型在四种典型机器人操作任务和两种机器人本体上的908次成功执行,适配并验证了八项不确定性指标与五项质量指标。通过领域专家人工标注任务质量,建立人类评判基准。结果表明,若干指标与专家评分呈中到强相关,具备评估执行质量与模型置信度的潜力。此外,部分指标能有效区分失败任务中高质量、中等质量与低质量执行,为缺乏成功标签时的评估提供支持。研究挑战了仅以成功率评价机器人的局限性,为实时监控与自适应优化提供了新路径。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA)-enabled robots integrate visual perception, natural language understanding, and action planning to interpret their environment, comprehend instructions, and perform embodied tasks autonomously. Such robots are typically evaluated through task success rates, i.e., whether a robot performs its intended task, which are commonly used as test oracles for evaluating such robots. Such an evaluation fails to capture the quality of task execution and the robot's confidence in its decisions. In this paper, we adapt eight uncertainty metrics and five quality metrics specifically designed for VLA-enabled robotic manipulation tasks. We assess their effectiveness through a large-scale empirical study involving 908 successful task executions from three state-of-the-art VLA models across four representative robotic manipulation tasks and two robot embodiments. Human domain experts manually labeled task quality, enabling us to analyze the correlation between our proposed metrics and expert judgments, serving as a human oracle for testing such robots. The results reveal that several metrics show moderate to strong correlation with human assessments, highlighting their utility for evaluating task quality and model confidence. Furthermore, we found that some metrics can discriminate between high-, medium-, and low-quality executions from unsuccessful tasks, which is useful when test oracles are absent. Our findings challenge the adequacy of current evaluation practices that rely solely on binary success rates and pave the way for improved real-time monitoring and adaptive enhancement of VLA-enabled robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。