arXiv:2510.08884cs.RO2025-10被引 2

用模型预测提升抓取任务表现,兼顾精度与泛化能力

Model-Based Lookahead Reinforcement Learning for in-hand manipulation

  • 结合模型预测与强化学习,通过动态模型评估轨迹优化控制
  • 在多种物体属性变化下仍保持性能提升,平均奖励显著改善
  • 适合需要高精度和强泛化的机器人抓取场景

手部操作作为复杂灵巧任务的代表,在机器人领域仍面临巨大挑战,需同时处理复杂的动态系统并精确操控各类物体。本文将一种混合强化学习框架应用于手部操作任务,验证其可有效提升任务表现。该框架融合了无模型与基于模型的强化学习思想,利用动态模型和价值函数指导训练好的策略进行轨迹评估,类似模型预测控制(MPC)机制。实验在全驱动与欠驱动的仿真机械手平台上展开,测试了不同物体的操控任务。结果表明,在具备高平均奖励的策略与准确动态模型的前提下,该混合框架在多数测试案例中均能提升手部操作性能,即使在物体密度、尺寸等属性发生变化时也具备良好泛化能力。然而,由于轨迹评估计算复杂度较高,整体计算开销有所增加。

原文摘要 · Abstract (English)

In-Hand Manipulation, as many other dexterous tasks, remains a difficult challenge in robotics by combining complex dynamic systems with the capability to control and manoeuvre various objects using its actuators. This work presents the application of a previously developed hybrid Reinforcement Learning (RL) Framework to In-Hand Manipulation task, verifying that it is capable of improving the performance of the task. The model combines concepts of both Model-Free and Model-Based Reinforcement Learning, by guiding a trained policy with the help of a dynamic model and value-function through trajectory evaluation, as done in Model Predictive Control. This work evaluates the performance of the model by comparing it with the policy that will be guided. To fully explore this, various tests are performed using both fully-actuated and under-actuated simulated robotic hands to manipulate different objects for a given task. The performance of the model will also be tested for generalization tests, by changing the properties of the objects in which both the policy and dynamic model were trained, such as density and size, and additionally by guiding a trained policy in a certain object to perform the same task in a different one. The results of this work show that, given a policy with high average reward and an accurate dynamic model, the hybrid framework improves the performance of in-hand manipulation tasks for most test cases, even when the object properties are changed. However, this improvement comes at the expense of increasing the computational cost, due to the complexity of trajectory evaluation.

强化学习手部操作模型预测机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。