让机器人控制模型自己判断靠谱程度,提前发现危险。
ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control

- 用校准的置信度估计动作输出,反映真实可靠性
- 在仿真和真实机器人上均提升失败预警能力,降低灾难性错误
- 无需重训练即可增强现有视觉语言动作模型的安全性
视觉-语言-动作(VLA)模型已成为能够将视觉观测和自然语言指令映射为连续动作序列的通用机器人控制器。然而,现有VLA模型无法提供动作预测的校准置信度,限制了其在需预判不确定性和故障的真实场景中的可靠性。为此,我们提出ReconVLA,一种可生成不确定性引导与故障感知控制信号的可靠符合性模型。具体而言,我们的方法直接对预训练VLA策略的动作标记输出应用符合性预测,获得与执行质量及任务成功率相关联的校准不确定性估计。此外,我们将符合性预测扩展至机器人状态空间,以在故障发生前检测异常或不安全状态,提供简单而有效的故障检测机制,补充动作层面的不确定性。我们在多种操作任务的仿真与真实机器人实验中评估了ReconVLA。结果表明,符合性化动作预测能持续提升失败预见能力,减少灾难性错误,并在不重新训练或修改底层VLA的前提下提供校准的置信度测量。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have emerged as generalist robotic controllers capable of mapping visual observations and natural language instructions to continuous action sequences. However, VLAs provide no calibrated measure of confidence in their action predictions, thus limiting their reliability in real-world settings where uncertainty and failures must be anticipated. To address this problem we introduce ReconVLA, a reliable conformal model that produces uncertainty-guided and failure-aware control signals. Concretely, our approach applies conformal prediction directly to the action token outputs of pretrained VLA policies, yielding calibrated uncertainty estimates that correlate with execution quality and task success. Furthermore, we extend conformal prediction to the robot state space to detect outliers or unsafe states before failures occur, providing a simple yet effective failure detection mechanism that complements the action-level uncertainty. We evaluate ReconVLA in both simulation and real robot experiments across diverse manipulation tasks. Our results show that conformalized action predictions consistently improve failure anticipation, reduce catastrophic errors, and provide a calibrated measure of confidence without retraining or modifying the underlying VLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。