提出自评估框架MAE,让视觉语言动作模型自己判断动作可靠性。
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy

- 利用注意力熵构建马尔可夫链模型,捕捉内部信号差异
- 在4000集任务中表现超越现有方法,AUPR、AUROC等指标领先
- 适合需要无监督可靠性评估的机器人决策场景
视觉-语言-动作模型(VLAs)将视觉感知、语言指令与动作生成统一为端到端策略,但如何在无外部监督下自评估动作生成可靠性仍是难题。现有方法依赖专家标注或仅从输出统计估算不确定性,忽视内部信号。本文发现,不同架构的VLAs在成功与失败任务中,视觉模态的熵值存在一致差异。尽管动作生成结构各异,但它们共享一个由视觉感知、语言指令和状态输入共同驱动的隐式动作生成抽象,我们将其建模为条件生成马尔可夫链。基于此,提出MAE(Markov Attention Entropy)自评估框架,直接将内部注意力信号转化为适配架构的可靠性评分,并构建LIBERO-Reflect基准,包含2000个标准任务与2000个挑战任务,覆盖四个子集。跨多种异构VLA架构与场景的实验表明,MAE在AUPR、AUROC和FPR@95上持续优于当前最优基线。进一步实例化为FabriMAE,实现无需验证器的测试时动作选择,在LIBERO-Plus上提升PI系列模型鲁棒性,且观测运行开销极小。
原文摘要 · Abstract (English)
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and state input, which we formulate as a Conditional Generative Markov Chain. Based on this formulation, we propose MAE (Markov Attention Entropy), a self-evaluation framework that directly converts internal attention signals into architecture-aware reliability scores, and introduce LIBERO-Reflect, a 4,000-episode benchmark combining 2,000 standard episodes and 2,000 challenging episodes across four subsets. Extensive experiments across heterogeneous VLA architectures and diverse scenarios show that MAE consistently outperforms state-of-the-art baselines on AUPR, AUROC, and FPR@95. We further instantiate FabriMAE for verifier-free test-time action selection, showing that MAE-guided multiple sampling improves PI-family robustness on LIBERO-Plus with small observed runtime overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。