arXiv:2409.14412cs.LGcs.AI2024-09

用不完美的仿真数据提升离线强化学习性能,解决真实环境难获取问题

COSBO: Conservative Offline Simulation-Based Policy Optimization

  • 融合真实数据与有偏差的仿真数据训练策略
  • 在复杂动态场景下超越CQL、MOPO等先进方法
  • 适合无法直接交互真实世界的机器人控制任务

离线强化学习可在真实部署数据上训练模型,但仅能选择训练数据中已有的行为组合。而仿真环境虽可替代真实数据,却存在仿真到现实的差距,引入偏差。为此,我们提出一种结合不完美仿真环境与目标环境数据的方法,用于训练离线强化学习策略。实验表明,该方法在多样且具有挑战性的动态场景下,显著优于CQL、MOPO和COMBO等前沿方法,并在多种实验条件下表现稳健。结果表明,即使存在仿真到现实的差距,利用仿真生成数据仍可有效提升离线策略学习效果,尤其在无法直接与真实世界交互的情况下。

原文摘要 · Abstract (English)

Offline reinforcement learning allows training reinforcement learning models on data from live deployments. However, it is limited to choosing the best combination of behaviors present in the training data. In contrast, simulation environments attempting to replicate the live environment can be used instead of the live data, yet this approach is limited by the simulation-to-reality gap, resulting in a bias. In an attempt to get the best of both worlds, we propose a method that combines an imperfect simulation environment with data from the target environment, to train an offline reinforcement learning policy. Our experiments demonstrate that the proposed method outperforms state-of-the-art approaches CQL, MOPO, and COMBO, especially in scenarios with diverse and challenging dynamics, and demonstrates robust behavior across a variety of experimental conditions. The results highlight that using simulator-generated data can effectively enhance offline policy learning despite the sim-to-real gap, when direct interaction with the real-world is not possible.

离线强化学习仿真增强策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。