arXiv:2607.13818cs.RO2026-07

让机器人通过智能决策恢复任务执行,提升操作鲁棒性。

Learning Robust Execution in Robotic Manipulation with Agentic Reinforcement Learning

论文配图:Learning Robust Execution in Robotic Manipulation with Agentic Reinforcement Learning
图 1 · 摘自论文原文
  • 用两个运行时指标评估执行质量,指导恢复决策
  • 在LIBERO上成功率达标准设置13.7%、扰动下39.2%提升
  • 适合需要长期稳定执行的复杂机械臂任务

机器人操作因不确定性、长时程执行和误差累积面临根本挑战,易导致执行失稳与任务失败。尽管近期视觉-语言-动作(VLA)模型具备强泛化能力,但通常缺乏对执行稳定性评估及偏离后恢复的显式机制。本文提出:(1) 两种互补的运行时执行质量评估指标;(2) 一种基于智能体的强化学习框架,通过高层决策而非直接学习底层动作来恢复有效执行。该框架中,智能体基于近期执行历史,从少量执行模式中选择并调控过程。当执行退化时,自动触发相应恢复机制,使机器人回到已访问的正常状态,从而继续任务。在LIBERO基准上评估,标准设置下成功率提升最高达13.7%,扰动设置下最高提升39.2%,显著增强执行鲁棒性。

原文摘要 · Abstract (English)

Robotic manipulation poses fundamental challenges due to uncertainty, long-horizon execution, and compounding errors, which can easily destabilize execution and lead to task failure. Although recent vision-language-action (VLA) models exhibit strong generalization, they typically lack explicit mechanisms to assess execution stability and to recover when execution deviates from its nominal behavior. In this paper, we propose: (1) two complementary metrics to assess execution quality at runtime, and (2) an agentic reinforcement learning framework that learns to restore effective execution through high-level decision-making rather than directly learning low-level actions. In this framework, an agentic policy reasons over recent execution history and selects among a small set of execution modes to regulate the execution process. Under execution degradation, it triggers appropriate recovery mechanisms to restore the robot to previously visited nominal states, enabling the task to continue. We evaluate the proposed method on the LIBERO benchmark, achieving up to a 13.7% improvement in success rate under standard settings and up to a 39.2% improvement under disturbance settings, demonstrating substantially enhanced execution robustness.

机器人操作强化学习鲁棒性智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。