arXiv:2502.13187cs.LGcs.AI2025-02综述被引 61

系统梳理强化学习中从仿真到现实的迁移方法与挑战。

A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models

  • 按马尔可夫决策过程四要素构建仿真到现实迁移的分类框架
  • 涵盖经典至前沿方法,包括基础模型赋能的新技术
  • 提供可复现评估方案,适合机器人、交通等领域研究者

深度强化学习在机器人、交通、推荐系统等多个领域被证明能有效解决决策问题,其通过与环境交互并基于经验更新策略。然而,由于真实世界数据有限且错误动作后果严重,策略学习主要局限于模拟器内进行。这一做法虽保障了训练安全,却不可避免地引入仿真到现实的差距,导致部署时性能下降甚至执行风险。为应对该问题,各领域已提出多种技术方案,尤其在大模型等新兴技术背景下,基础模型为仿真到现实迁移带来新思路。本文首次系统性地从马尔可夫决策过程的核心要素(状态、动作、转移、奖励)出发,构建仿真到现实迁移的正式分类体系。基于此框架,全面综述从经典到最先进方法,涵盖基础模型驱动的先进技术,并分析不同领域中需关注的特殊性。同时,总结可复现的评估流程及公开代码或基准。最后,指出当前挑战与未来机遇,以推动该方向发展。我们持续维护一个开源仓库,收录最新仿真到现实研究工作,供领域研究人员参考。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (RL) has been explored and verified to be effective in solving decision-making tasks in various domains, such as robotics, transportation, recommender systems, etc. It learns from the interaction with environments and updates the policy using the collected experience. However, due to the limited real-world data and unbearable consequences of taking detrimental actions, the learning of RL policy is mainly restricted within the simulators. This practice guarantees safety in learning but introduces an inevitable sim-to-real gap in terms of deployment, thus causing degraded performance and risks in execution. There are attempts to solve the sim-to-real problems from different domains with various techniques, especially in the era with emerging techniques such as large foundations or language models that have cast light on the sim-to-real. This survey paper, to the best of our knowledge, is the first taxonomy that formally frames the sim-to-real techniques from key elements of the Markov Decision Process (State, Action, Transition, and Reward). Based on the framework, we cover comprehensive literature from the classic to the most advanced methods including the sim-to-real techniques empowered by foundation models, and we also discuss the specialties that are worth attention in different domains of sim-to-real problems. Then we summarize the formal evaluation process of sim-to-real performance with accessible code or benchmarks. The challenges and opportunities are also presented to encourage future exploration of this direction. We are actively maintaining a repository to include the most up-to-date sim-to-real research work to help domain researchers.

强化学习仿真到现实基础模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。