让视觉语言动作模型根据任务难易自动选择执行、思考或放弃,提升效率与安全。
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
- 根据感知状态复杂度动态决策:执行、思考或中止
- 在真实机器人上达87.5%的F1分数,仅用5%数据仍保持83%性能
- 首次实现对分布外场景的不确定性估计,适合高可靠性应用
当前视觉-语言-动作(VLA)模型研究主要通过推理技术提升泛化能力,但显著增加计算开销和推理延迟。且这些机制常被无差别使用,导致简单任务资源浪费,无法有效估计不确定性以避免分布外场景下的灾难性失败。受人类认知启发,我们提出一种自适应框架,基于感知状态复杂度动态路由VLA执行流程。将视觉-语言主干转化为主动检测工具,通过将隐向量投影到参数化与非参数化估计器中,实现已知任务的即时执行(Act)、模糊场景的推理(Think),以及物理或语义异常时的主动中止(Abstain)。我们发现,对融合视觉-语言嵌入拟合的高斯混合模型能提供最可靠的任务复杂度信号,结合视觉新颖性、指令上下文与跨模态兼容性。在LIBERO和LIBERO-PRO基准及真实机器人上的评估显示,融合配置在两种VLA主干(SmolVLA和π₀)上达到最高87.5% F1分数,仅用5%训练数据仍维持83%性能,超越现有最先进故障检测器。
原文摘要 · Abstract (English)
Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effective, these improvements increase computational complexity and inference latency. Furthermore, these mechanisms are typically applied indiscriminately, wasting resources on trivial tasks while failing to provide the uncertainty estimation necessary to prevent catastrophic failure on out-of-distribution scenarios. Inspired by human cognition, we propose an adaptive framework that dynamically routes VLA execution based on the complexity of the perceived state. Our approach transforms the VLA's vision-language backbone into an active detection tool by projecting latent embeddings into a set of parametric and non-parametric estimators. This allows the system to execute known tasks immediately (Act), reason about ambiguous scenarios (Think), and preemptively halt execution when encountering physical or semantic anomalies (Abstain). We find that a Gaussian Mixture Model fitted to fused vision-language embeddings provides the most reliable task-complexity signal, combining visual novelty with instruction context and cross-modal compatibility. Evaluated on the LIBERO and LIBERO-PRO benchmarks as well as on a real robot, our fused configuration achieves up to 87.5% F1-score across two VLA backbones (SmolVLA and $π_0$), retains 83% with as little as 5% of training data, and surpasses state-of-the-art failure detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。