arXiv:2605.01194cs.RO2026-05被引 14

让视觉语言动作模型在复杂任务中自动思考,减少错误决策。

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model

论文配图:VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
图 1 · 摘自论文原文
  • 用不确定度触发动态思考模式,替代快速直觉反应。
  • 通过成对比较选出最优动作,失败率比当前最好模型降低50%以上。
  • 无需人工标注,自动构建训练数据,适合需要高可靠性的机器人任务。

视觉-语言-动作(VLA)模型在具身操作中展现出强大能力与泛化性,但其决策依赖快速直觉过程,缺乏深度思考。这在复杂或模糊场景下常导致次优甚至灾难性动作。本文提出VLA-ATTC框架,赋予VLA模型自适应测试时计算(TTC)能力。该框架采用基于不确定度的“认知离合器”,在必要时动态切换至TTC思辨阶段。在该阶段,新颖的相对动作评价(RAC)模型通过成对比较从生成候选动作中识别最优解,替代不稳定的绝对值估计,显著简化学习目标。此外,引入高效采样策略以分摊计算开销,并设计自动化数据流水线,无须人工标注即可构建偏好对。在LIBERO-LONG基准上,VLA-ATTC使当前最优模型PI0.5的失败率降低超50%。代码与权重将全部开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities and generalization in embodied manipulation. However, their decision-making relies on a fast, instinctive process that lacks deliberation. This strategy often leads to suboptimal or catastrophic actions when facing complex or ambiguous scenarios that require greater consideration. In this paper, we introduce \textbf{VLA-ATTC}, a framework that endows VLA models with adaptive test-time compute (TTC). VLA-ATTC employs an uncertainty-based ``cognitive clutch'' to dynamically transition from reflexive execution to a TTC deliberation phase when necessary. During TTC phase, a novel \textbf{Relative Action Critic} (RAC) model identifies the optimal action from generated candidates via pairwise comparisons. This relative mechanism replaces unstable absolute value estimation, significantly simplifying the learning objective. Furthermore, we introduce an efficient sampling strategy to amortize computational costs and an automated data pipeline that curates preference pairs without manual annotation. On the LIBERO-LONG benchmark, VLA-ATTC reduces the failure rate of the SOTA model PI0.5 by over 50\%. We will open-source all the code and weights.

机器人自适应计算强化学习视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。