让视觉语言动作模型学会准确判断自己能否成功,提升机器人可信度。
Confidence Calibration in Vision-Language-Action Models
- 首次研究视觉语言动作模型的置信度校准问题。
- 发现任务成功率与置信度误差相关,且随时间变化。
- 提出轻量级方法改善置信度不准问题,适合部署在真实机器人上。
可信的机器人行为不仅需要高任务成功率,还需能可靠地量化成功可能性。为此,我们首次对视觉-语言-动作(VLA)基础模型的置信度校准进行了系统研究,这类模型将视觉观测和自然语言指令映射为低层机器人运动命令。我们建立了VLA的置信度基准,分析了任务成功率与校准误差的关系,并考察了校准随时间的演变规律;同时引入两种轻量级修正技术:提示集成与逐动作Platt缩放。本研究旨在初步构建使VLA兼具高性能与高可信度所必需的工具与概念理解,通过可靠的不确定性量化实现可信自主决策。
原文摘要 · Abstract (English)
Trustworthy robot behavior requires not only high levels of task success but also that the robot can reliably quantify how likely it is to succeed. To this end, we present a first-of-its-kind study of confidence calibration in vision-language-action (VLA) foundation models, which map visual observations and natural language instructions to low-level robot motor commands. We establish a confidence baseline for VLAs, examine how task success relates to calibration error and how calibration evolves over time, and introduce two lightweight techniques to remedy the miscalibration we observe: prompt ensembles and action-wise Platt scaling. Our aim in this study is to begin to develop the tools and conceptual understanding necessary to render VLAs both highly performant and highly trustworthy via reliable uncertainty quantification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。