通过对比扰动训练提升视觉语言动作策略的故障检测鲁棒性
SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

- 用对比样本扰动增强隐藏状态探测器的训练与校准
- 在真实和仿真环境中故障检测ROC-AUC显著优于基线方法
- 视觉与语言双重扰动效果最佳,适合部署时分布偏移场景
视觉-语言-动作策略在部署时遭遇遮挡、干扰物、光照变化、新物体、初始状态改变或指令重述等分布外情况时常失效。基于隐藏状态的风险探测器结合功能型共形预测可检测执行失败,但其可靠性依赖于与部署环境匹配的校准数据。本文提出SAFECAST,利用对比集扰动改进隐藏状态探测器的训练与校准,以应对部署时的分布偏移。在真实世界DROID和模拟环境LIBERO中,SAFECAST对多个VLM骨干模型均显著提升故障检测的ROC-AUC表现。进一步发现,同时使用视觉与语言对比扰动时收益最大;且相比仅使用真实回放数据,基于仿真到现实的校准能生成更优探测器。
原文摘要 · Abstract (English)
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。