通过扰动权重检测视觉语言动作模型的不确定性,提升机器人在未知场景下的故障预警能力。
Perturbation-Based Epistemic Uncertainty for Failure Detection in Vision-Language-Action Models

- 用低秩权重扰动生成多个预测,通过结果不一致度衡量模型自身不确定性
- 在LIBERO-PRO数据集上,平均AUROC和平衡准确率均优于现有方法
- 无需训练即可部署,适合真实机器人在未知环境中的故障检测
视觉-语言-动作(VLA)模型在机器人操作中表现优异,但可靠的不确定性量化仍具挑战性,尤其在分布偏移情况下。与自回归策略不同,多数现代VLA模型通过回归或流式生成连续动作,无法直接提供预测概率。此外,随机动作采样仅反映固定模型下的生成变异性,而分布偏移下的故障检测更需捕捉模型本身的不确定性。受贝叶斯局部模型变化视角启发,本文提出无需训练的扰动型故障检测(PFD)框架,通过向选定Transformer权重矩阵注入随机低秩扰动,基于扰动后动作预测的分歧度估计后验不确定性。在LIBERO-PRO上的实验表明,PFD在所评估方法中达到最高平均AUROC与平衡准确率,且在各类分布偏移下持续优于随机动作采样。真实机器人实验进一步验证了其在未见物体转移情境下的有效故障检测能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but reliable uncertainty quantification remains challenging, particularly under distribution shift. Unlike autoregressive policies, many modern VLA models generate continuous actions through regression or flow-based generation, where explicit predictive probabilities are unavailable. Moreover, stochastic action sampling primarily captures action-generation variability under a fixed model, while failure detection under distribution shift can benefit from capturing uncertainty in the model itself. Motivated by Bayesian perspectives on local model variations, we propose perturbation-based failure detection (PFD), a training-free framework for estimating epistemic uncertainty in VLA models through low-rank weight perturbations. Specifically, we inject random low-rank perturbations into selected transformer weight matrices and estimate epistemic uncertainty from disagreement across perturbed action predictions. Experiments on LIBERO-PRO show that PFD achieves the highest average AUROC and balanced accuracy among the evaluated methods while consistently outperforming stochastic action sampling across distribution shifts. Real-world robot experiments further demonstrate that PFD provides a competitive failure-detection signal under an unseen object shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。