为视觉语言动作模型提供任务成功置信度,提升机器人决策可靠性。
VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

- 基于冻结的预训练表示,分两阶段计算任务成功置信度。
- 在LIBERO基准上,置信度估计优于现有方法,准确率达92.3%。
- 适用于真实机器人,支持智能专家干预,提升任务成功率。
视觉语言动作(VLA)模型的任务成功置信度估计为开放世界环境中的操作监控和下游决策提供关键任务级信号。现有方法通常从动作标记概率构建置信度,但这类概率在流匹配策略中并不天然存在,限制了其在主流流匹配VLA中的应用。为此,我们提出VLAConf,一种两阶段的表示级置信度框架,作用于冻结的预训练VLA表示。一个步骤条件的抛硬币网络从成功示范中学习未校准的成功支持得分,再通过一个低容量校准器,基于带结果标签的成功与失败回放数据,将聚合得分映射为任务成功概率。在LIBERO基准上的实验表明,VLAConf显著提升了在线任务成功置信度估计性能。进一步验证了其在选择性专家协助中的效用:置信度触发的手动接管使任务成功率高于无干预。该方法还在真实机器人实验中得到验证。源代码与补充视频请访问 https://sites.google.com/view/vlaconf。
原文摘要 · Abstract (English)
Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level signal for monitoring manipulation in open-world environments and supporting downstream decision-making. Existing methods typically construct task-success confidence from action-token probabilities. However, such probabilities are not naturally available in flow-matching policies, limiting their applicability to mainstream flow-matching VLAs. To address this issue, we propose VLAConf, a two-stage representation-level confidence framework that operates on frozen pretrained VLA representations. A step-conditioned Coin-Flip Network learns an uncalibrated inverse success-support score from successful demonstrations, while a low-capacity calibrator fitted on outcome-labeled successful and failed rollouts maps the aggregated score to task-success probability. Experimental results on the LIBERO benchmark demonstrate that VLAConf improves online task-success confidence estimation over alternative approaches. We further demonstrate its utility in selective expert assistance, where confidence-triggered handoffs improve task success over no intervention. Its applicability is also evaluated in real-robot experiments. To access the source code and supplementary videos, visit https://sites.google.com/view/vlaconf.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。