让视觉语言模型直接学会用潜在动作条件控制机器人,提升操控精度。
CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models

- 在视觉语言模型中引入轻量级潜在动作接口,实现原生动作条件化。
- 在LIBERO数据集上达到98.3%成功率,LIBERO-Plus达89.5%。
- 适合需要高精度连续动作控制的通用机器人任务研究者。
视觉-语言-动作(VLA)模型已成为通用机器人操作的有前景范式,利用视觉语言表征来引导连续动作生成。然而,这些表征并未显式优化用于动作条件化,导致动作专家需自行弥合多模态理解与精确运动控制之间的差距。近期动作推理方法引入额外模块生成显式动作规划或动作空间推理信号,虽证明了动作级引导的有效性,但通常需要独立的动作生成框架。本文提出CAC-VLA,一种上下文门控动作条件化框架,将轻量级潜在动作接口直接学习于视觉语言模型(VLM)内部。不同于生成可执行轨迹,CAC-VLA训练VLM预测粗到细的潜在动作,即从未来动作片段编码出的结构化表示,并通过上下文门控机制自适应地将其用于条件化动作专家。该方法实现了VLM原生的动作条件化,同时校准了潜在动作引导对专家动作生成的影响。在LIBERO和LIBERO-Plus上的实验表明,CAC-VLA表现优异,平均成功率达98.3%(LIBERO)和89.5%(LIBERO-Plus),表明上下文门控的潜在动作条件化是连续专家控制的有效接口。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have become a promising paradigm for generalist robot manipulation, where visual-language representations are used to condition continuous action generation. However, these representations are not explicitly optimized for action conditioning, leaving the action expert to bridge the gap between multimodal understanding and precise motor control. Recent action-reasoning methods introduce additional modules to generate explicit action plans or action-space reasoning signals, demonstrating the benefit of action-level guidance but often requiring separate action-generation frameworks. We propose CAC-VLA, a Context-Gated Action Conditioning framework that learns a lightweight latent-action interface directly within the VLM. Instead of generating executable trajectories, CAC-VLA trains the VLM to predict coarse-to-fine latent actions, which are structured representations encoded from future action segments, and adaptively leverages them to condition the action expert via a context gate. This enables VLM-native action conditioning while calibrating the influence of latent-action guidance on expert action generation. Experiments on LIBERO and LIBERO-Plus demonstrate the effectiveness of CAC-VLA, achieving 98.3% average success rate on LIBERO and 89.5% LIBERO-Plus, suggesting that context-gated latent-action conditioning is an effective interface for continuous expert control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。