通过对齐视觉动作表示,提升机器人从图像和语言中学习动作的能力。
LARA: Latent Action Representation Alignment for Vision-Language-Action Models

- 联合优化视觉-语言-动作模型与潜在动作模型,实现双向增强。
- 在模拟与真实场景中分别提升10%、5%、15%的性能表现。
- 适合需要提升动作预测准确性的机器人学习研究者使用。
视觉-语言-动作(VLA)模型使机器人能够直接根据观测和语言指令预测动作,但其性能依赖大规模高质量数据,且受限于真实机器人动作数据集稀缺。为利用大量未标注的人类视频进行学习,潜在动作模型(LAM)从视觉动态中学习潜在动作表示,为VLA学习提供额外监督。然而,LAM与VLA通常独立训练,导致LAM在VLA训练时缺乏实际动作语境,而VLA受制于冻结的LAM表示。为此,我们提出潜行动作表示对齐(LARA)框架,通过表示对齐联合优化LAM与VLA。该方法实现相互增益:LAM结合动作轨迹学习,避免虚假视觉变化;VLA则借助LAM内部学习的前向动力学,减少功能无效轨迹的幻觉。我们在3个仿真和1个精心设计的真实机器人操作基准上验证了LARA的通用性与有效性,分别实现约10%、5%、15%的性能提升。
原文摘要 · Abstract (English)
Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。