让机器人视觉语言模型更智能地省算力,按动作上下文动态调整计算。
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
- 根据视觉、语言和动作状态动态调整模型计算,减少冗余
- 速度提升1.79倍,计算量降至原模型的29.4%,成功率相当
- 适合需要低延迟部署的机器人操作任务
视觉-语言-动作(VLA)模型在机器人操作中表现优异,但闭环部署受限于每步重复运行大型视觉-语言主干带来的高延迟与高算力开销。我们发现VLA推理在时间、空间和深度维度上存在结构性冗余,而现有效率方法普遍忽略动作上下文,尽管其在具身任务中至关重要。为此,提出面向动作上下文感知的自适应计算框架AC^2-VLA,该框架基于当前视觉观测、语言指令和先前动作状态,统一实现跨时间步的认知复用、标记剪枝与组件选择性执行。为训练自适应策略,引入一种动作引导的自蒸馏方案,在保留密集VLA策略行为的同时实现结构化稀疏化,且具备跨任务与场景的迁移能力。在多个机器人操作基准测试中,AC^2-VLA在保持相当任务成功率的前提下,最高实现1.79倍加速,仅需原模型29.4%的浮点运算量。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have demonstrated strong performance in robotic manipulation, yet their closed-loop deployment is hindered by the high latency and compute cost of repeatedly running large vision-language backbones at every timestep. We observe that VLA inference exhibits structured redundancies across temporal, spatial, and depth dimensions, and that most existing efficiency methods ignore action context, despite its central role in embodied tasks. To address this gap, we propose Action-Context-aware Adaptive Computation for VLA models (AC^2-VLA), a unified framework that conditions computation on current visual observations, language instructions, and previous action states. Based on this action-centric context, AC^2-VLA adaptively performs cognition reuse across timesteps, token pruning, and selective execution of model components within a unified mechanism. To train the adaptive policy, we introduce an action-guided self-distillation scheme that preserves the behavior of the dense VLA policy while enabling structured sparsification that transfers across tasks and settings. Extensive experiments on robotic manipulation benchmarks show that AC^2-VLA achieves up to a 1.79\times speedup while reducing FLOPs to 29.4% of the dense baseline, with comparable task success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。