让机器人通过动态推理实现自适应操作,成功率接近100%。
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning

- 在强化学习中融合思维链式推理,动态调节思考深度。
- 在LIBERO基准上达99.9%成功率,仅需一次监督预热。
- 适合需要高适应性与真实世界部署的机器人任务研究者。
机器人基础模型需在动态环境中对复杂视觉场景进行推理以执行自适应动作。尽管近期基于潜在空间推理的视觉-语言-动作(VLA)模型已能捕捉精细物理动态,但主要局限于静态模仿学习,严重限制其适应性和泛化能力。本文提出LaST-R1,一种新型强化学习(RL)后训练框架,旨在有效利用“先推理再行动”的策略。我们设计了潜在于动作策略优化(LAPO)算法,联合优化潜在推理过程与动作生成。通过将潜在思维链(CoT)推理直接嵌入强化学习优化循环,LAPO激发深层物理世界建模,从而提升交互环境中的稳健执行。此外,引入自适应潜在思维链机制,使策略可根据环境状态动态调节推理深度。实验表明,LaST-R1在LIBERO基准上仅需一次监督预热即达到99.9%平均成功率,显著提升收敛速度与性能。在真实部署中,相比最优监督微调方法,平均提升达22.5%,涵盖单臂与双臂共四项复杂任务。最终,LaST-R1在仿真与真实环境间展现出强泛化能力。
原文摘要 · Abstract (English)
Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language-Action (VLA) models have demonstrated the capability to capture fine-grained physical dynamics, they remain predominantly confined to static imitation learning, severely limiting their adaptability and generalization. In this paper, we present LaST-R1, a novel reinforcement learning (RL) post-training framework designed to effectively harness "latent reasoning-before-acting" policies. Specifically, we propose Latent-to-Action Policy Optimization (LAPO), a core RL algorithm that jointly optimizes the latent reasoning process and the action generation. By explicitly embedding latent Chain-of-Thought (CoT) reasoning directly within the RL optimization loop, LAPO stimulates profound physical world modeling, which in turn drives robust execution in interactive environments. Furthermore, an adaptive latent CoT mechanism is introduced, allowing the policy to dynamically modulate its reasoning horizon based on diverse environment states. Experiments show that LaST-R1 achieves a near-perfect 99.9% average success rate on the LIBERO benchmark with only one-shot supervised warm-up, significantly improving convergence speed and performance over prior state-of-the-art (SOTA) methods. In real-world deployments, LaST-R1 yields up to a 22.5% average improvement over SOTA supervised fine-tuning approach across four complex tasks, including both single-arm and dual-arm settings. Finally, LaST-R1 demonstrates strong generalization across simulated and real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。