arXiv:2604.01142cs.ROcs.LG2026-04被引 2

用强化学习加边界极值搜索提升机器人操作在变化环境中的鲁棒性。

Deep Reinforcement Learning for Robotic Manipulation under Distribution Shift with Bounded Extremum Seeking

  • 结合DDPG强化学习与有界极值搜索,实现快速操作与环境适应。
  • 在目标变化、摩擦不均等分布外场景下仍保持稳定性能。
  • 适合需要高鲁棒性的复杂抓取与推移任务研究者参考。

强化学习在机器人操作中表现优异,但当测试条件偏离训练分布时性能常显著下降。这一问题在接触丰富的任务(如推移和抓放)中尤为突出,因目标变化、接触条件或机器人动力学改变可能导致推理时系统分布外。本文提出一种混合控制器,将深度确定性策略梯度(DDPG)策略在标准条件下训练后,部署时结合有界极值搜索(bounded extremum seeking)。RL策略提供快速操作行为,而有界ES确保系统在运行条件偏离训练数据时具备时间变化鲁棒性。该控制器在多种分布外设置下进行评估,包括时变目标和空间变化的摩擦区域。

原文摘要 · Abstract (English)

Reinforcement learning has shown strong performance in robotic manipulation, but learned policies often degrade in performance when test conditions differ from the training distribution. This limitation is especially important in contact-rich tasks such as pushing and pick-and-place, where changes in goals, contact conditions, or robot dynamics can drive the system out-of-distribution at inference time. In this paper, we investigate a hybrid controller that combines reinforcement learning with bounded extremum seeking to improve robustness under such conditions. In the proposed approach, deep deterministic policy gradient (DDPG) policies are trained under standard conditions on the robotic pushing and pick-and-place tasks, and are then combined with bounded ES during deployment. The RL policy provides fast manipulation behavior, while bounded ES ensures robustness of the overall controller to time variations when operating conditions depart from those seen during training. The resulting controller is evaluated under several out-of-distribution settings, including time-varying goals and spatially varying friction patches.

强化学习机器人操控鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。