综述如何用迁移与逆强化学习提升强化学习的样本效率和泛化能力。
Towards Sample-Efficiency and Generalization of Transfer and Inverse Reinforcement Learning: A Comprehensive Literature Review
- 通过迁移学习实现跨领域知识高效转移,结合人类反馈与仿真到真实场景策略。
- 逆强化学习采用少量经验过渡训练,支持多智能体与多意图问题建模。
- 适合关注强化学习效率与可扩展性的研究人员参考。
强化学习(RL)是机器学习的一个分支,主要解决智能体与环境交互中的序列决策问题,通过环境奖励来优化行为。然而,该范式因需大量数据收集而存在样本效率低、泛化困难的问题。此外,设计能权衡多种需求的显式奖励函数也极为耗时。近年来,迁移与逆强化学习(T-IRL)被用于缓解这些问题。本文系统综述了通过T-IRL实现强化学习算法的样本效率与泛化能力提升的研究进展。在简要介绍强化学习后,本文阐述了基础的T-IRL方法,并全面回顾了各领域的最新进展。研究发现,多数近期工作借助人机协同与仿真实验到真实世界(sim-to-real)策略,在迁移学习框架下实现了从源域到目标域的知识高效迁移。在逆强化学习中,研究重点集中在仅需少量经验轨迹的训练方案,以及将框架拓展至多智能体与多意图问题。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is a sub-domain of machine learning, mainly concerned with solving sequential decision-making problems by a learning agent that interacts with the decision environment to improve its behavior through the reward it receives from the environment. This learning paradigm is, however, well-known for being time-consuming due to the necessity of collecting a large amount of data, making RL suffer from sample inefficiency and difficult generalization. Furthermore, the construction of an explicit reward function that accounts for the trade-off between multiple desiderata of a decision problem is often a laborious task. These challenges have been recently addressed utilizing transfer and inverse reinforcement learning (T-IRL). In this regard, this paper is devoted to a comprehensive review of realizing the sample efficiency and generalization of RL algorithms through T-IRL. Following a brief introduction to RL, the fundamental T-IRL methods are presented and the most recent advancements in each research field have been extensively reviewed. Our findings denote that a majority of recent research works have dealt with the aforementioned challenges by utilizing human-in-the-loop and sim-to-real strategies for the efficient transfer of knowledge from source domains to the target domain under the transfer learning scheme. Under the IRL structure, training schemes that require a low number of experience transitions and extension of such frameworks to multi-agent and multi-intention problems have been the priority of researchers in recent years.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。