让机器人学会长时间按语言指令完成复杂操作
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- 用多视角视觉+自身状态构建可行动作空间
- 在长序列任务上性能超越现有最佳方法
- 适合需要连续决策的机器人实际应用
长时程、语言引导的机器人操作能力依赖于历史信息利用和连贯动作序列生成,但现有视觉-语言-动作(VLA)模型常忽视此点。为此,我们提出LoLA(长时程潜在动作学习)框架,融合长期多视角观测与机器人本体感知,实现多步推理与动作生成。首先通过视觉-语言模型编码历史序列与多视角信息;再引入状态感知潜在重表示模块,将视觉输入与语言指令映射至可执行的机器人运动空间。该模块通过可学习的‘具身锚定’潜在空间,显式地将视觉-语言表征与物理尺度对齐,区别于仅拼接本体感知与VL嵌入的传统方法。我们在多样化的机器人预训练数据集上训练LoLA,并在模拟基准(SIMPLER与LIBERO)及Franka和双臂Aloha机器人的真实任务上进行评估。结果表明,LoLA显著优于先前最先进方法(如pi0),尤其在长时程操作任务中表现突出。
原文摘要 · Abstract (English)
The capability of performing long-horizon, language-guided robotic manipulation tasks critically relies on leveraging historical information and generating coherent action sequences. However, such capabilities are often overlooked by existing Vision-Language-Action (VLA) models. To solve this challenge, we propose LoLA (Long Horizon Latent Action Learning), a framework designed for robot manipulation that integrates long-term multi-view observations and robot proprioception to enable multi-step reasoning and action generation. We first employ Vision-Language Models to encode rich contextual features from historical sequences and multi-view observations. We further introduces a key module, State-Aware Latent Re-representation, which transforms visual inputs and language commands into actionable robot motion space. Unlike existing VLA approaches that merely concatenate robot proprioception (e.g., joint angles) with VL embeddings, this module leverages such robot states to explicitly ground VL representations in physical scale through a learnable "embodiment-anchored" latent space. We trained LoLA on diverse robotic pre-training datasets and conducted extensive evaluations on simulation benchmarks (SIMPLER and LIBERO), as well as two real-world tasks on Franka and Bi-Manual Aloha robots. Results show that LoLA significantly outperforms prior state-of-the-art methods (e.g., pi0), particularly in long-horizon manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。