用信息论方法预测视觉语言动作模型的失败,跨环境通用且可解释。
Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

- 基于信息论构建三重信号,监测动作多样性、时间一致性与状态耦合。
- 在六个模型、三个环境中表现媲美最强基线,真实任务准确率达83%。
- 无需重训练即可跨模型、跨环境、跨仿真到现实迁移,适合安全部署场景。
视觉语言动作(VLA)模型广泛应用于各类任务,但其物理交互行为如同黑箱,可能造成不可逆损害,因此具备泛化性与可解释性的失败检测至关重要。我们发现成功与失败的执行过程具有系统性的信息论特征差异。基于此,将VLA控制建模为闭环信息流,并推导出三重信息论(Tri-Info)信号,用以捕捉动作是否保持多样性、时间上是否一致、是否与状态转移耦合。在六个VLA模型和三个基准环境中,Tri-Info在域内表现达到最强基线水平。更重要的是,该方法无需重新训练即可在不同架构、环境间以及仿真到现实的迁移中有效工作,在真实任务上实现83%的准确率,而先前检测器则退化至随机水平。Tri-Info不仅具备强大的跨域泛化失败检测能力,还能提供对失败原因的可解释诊断。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different information-theoretic signatures. Building on this, we formalize VLA control as a closed-loop information pipeline and derive the Triple Information-theoretic (Tri-Info) signals that capture whether actions remain diverse, temporally consistent, and coupled to state transitions. Across six VLA models and three benchmark environments, Tri-Info matches the strongest baselines in-domain. Moreover, Tri-Info transfers across architectures, environments, and the sim-to-real gap without retraining, reaching 83\% accuracy on real-world tasks where prior detectors collapse to chance. This establishes Tri-Info as a simple yet powerful method that not only detects failures with strong cross-domain generalization, but also delivers interpretable diagnostics of the underlying failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。