为流模型视觉语言动作系统提供不确定性量化,提升可靠性与适应性。
Uncertainty Quantification for Flow-Based Vision-Language-Action Models

- 通过小集合速度场分歧估算认知不确定性。
- 在LIBERO上验证其可预测失败,且比基线少需22%示范数据。
- 适合需要安全部署和高效微调的机器人任务场景。
视觉语言动作模型(VLAs)结合视觉语言主干与基于流匹配的大规模机器人数据训练的生成动作头,虽在机器人操作中表现优异,但缺乏预测置信度量化机制,无法识别不可靠动作。这在非平稳环境中构成重大隐患。为此,本文提出一种高效方法,利用小集成的速度场分歧(VFD)量化流匹配模型的认知不确定性。实验表明,该方法能有效用于部署时故障检测,并支持主动微调。我们提出SAVE框架,实现不确定性引导的主动多任务微调,显著减少对昂贵专家示范的需求。在LIBERO基准测试中,VFD获得更校准的不确定性估计,能准确预测下游性能,故障检测表现优越,且所需样本数比基线至少减少22%。结果证明,对流式VLAs进行不确定性量化,可同时提升故障感知与适应能力。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable. This presents a critical limitation for real-world deployment in non-stationary environments, where models inevitably encounter scenarios outside their pretraining distribution and may fail without warning. To address this, we derive an efficient method for quantifying epistemic uncertainty in flow-matching models by leveraging velocity-field disagreement (VFD) across a small ensemble. We successfully use this uncertainty estimate for failure detection during deployment and active fine-tuning of flow-based VLAs. To this end, we propose SAVE, a framework for uncertainty-guided active multitask fine-tuning that reduces the number of costly expert demonstrations required to adapt VLAs to new tasks. Through extensive experiments on the LIBERO benchmark, we demonstrate that VFD yields better-calibrated uncertainty estimates predictive of downstream performance, that VFD achieves strong performance in detecting failures, and that uncertainty-guided data acquisition with SAVE requires at least 22% fewer samples than baselines. In summary, our work shows that quantifying epistemic uncertainty in flow-based VLAs improves both failure awareness and adaptation. Project website: tum-lsy.github.io/uq_vla/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。