发现视觉分布偏移前的失败预警信号,验证了模型内部激活可预测崩溃。
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

- 通过分析OpenVLA内部激活,用线性探测器识别失败前兆。
- 层16探测器在遮挡任务中达AUROC 0.972,远超动作不一致的0.496。
- 适用于需要提前预警的机器人系统可靠性评估。
视觉分布偏移可能导致视觉-语言-动作策略在初始行为看似合理后突然失效。我们探究OpenVLA的内部激活是否包含失效前的信号。冻结策略后,每一步记录一个MLP激活,并训练两个线性监测器。遮挡使任务成功率从57%降至17%。在失败的匹配重置轨迹中,第16层逻辑探测器达到AUROC 0.972和AUPRC 0.352,而动作不一致仅达AUROC 0.496。未重新训练的遮挡训练探测器在相机抖动任务中仍达AUROC 0.689。然而校准检查显示,同一层16探测器在正常轨迹中平均每集触发3.32次预警。该对比表明,强回溯区分能力不等于实际运行中的安静预警行为。因拟合与评估共享任务、重置和随机种子,这些结果仅证明事后可分性,而非对独立样本的预测能力。
原文摘要 · Abstract (English)
Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with the steps before failure. We freeze the policy, record one MLP activation per LIBERO-10 step, and fit two linear monitors. Occlusion reduces task success from $57\%$ to $17\%$. Within failed matched-reset trajectories, a layer-16 logistic probe attains AUROC $0.972$ and AUPRC $0.352$, whereas action disagreement attains AUROC $0.496$. Without refitting, the occlusion-trained probe reaches AUROC $0.689$ on failed camera-jitter episodes. In a calibration check, however, the same layer-16 monitor averages 3.32 warning onsets per clean episode. This contrast shows that strong retrospective discrimination does not imply operationally quiet warning behavior. Because fitting and evaluation share tasks, resets, and seed, these results establish retrospective separability rather than prediction on independent episodes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。