arXiv:2409.07569cs.LGcs.AI2024-09综述被引 15

系统梳理约束强化学习的定义、方法与挑战,助力智能决策研究。

A Comprehensive Survey on Inverse Constrained Reinforcement Learning: Definitions, Progress and Challenges

  • 从专家示范中推断隐式约束,构建通用算法框架
  • 覆盖确定性/随机环境、少样本、多智能体等多种场景
  • 适合从事机器人、自动驾驶等安全敏感领域研究者

逆约束强化学习(ICRL)旨在基于专家示范数据推断其遵循的隐式约束。作为新兴研究方向,近年来受到广泛关注。本文对ICRL最新进展进行分类综述,为机器学习研究者和从业者提供全面参考,帮助初学者理解其定义、进展与关键挑战。首先形式化问题并提出跨场景的约束推断算法框架,涵盖确定性或随机环境、示范数据有限及多智能体情形。针对每类场景,阐明核心挑战并介绍基础方法。调查覆盖离散、虚拟与真实环境中的ICRL评估,并深入探讨自动驾驶、机器人控制、体育分析等应用。最后讨论未解关键问题,推动理论与工业应用间的衔接。相关文献可参见 https://github.com/Jasonxu1225/Awesome-Constraint-Inference-in-RL。

原文摘要 · Abstract (English)

Inverse Constrained Reinforcement Learning (ICRL) is the task of inferring the implicit constraints that expert agents adhere to, based on their demonstration data. As an emerging research topic, ICRL has received considerable attention in recent years. This article presents a categorical survey of the latest advances in ICRL. It serves as a comprehensive reference for machine learning researchers and practitioners, as well as starters seeking to comprehend the definitions, advancements, and important challenges in ICRL. We begin by formally defining the problem and outlining the algorithmic framework that facilitates constraint inference across various scenarios. These include deterministic or stochastic environments, environments with limited demonstrations, and multiple agents. For each context, we illustrate the critical challenges and introduce a series of fundamental methods to tackle these issues. This survey encompasses discrete, virtual, and realistic environments for evaluating ICRL agents. We also delve into the most pertinent applications of ICRL, such as autonomous driving, robot control, and sports analytics. To stimulate continuing research, we conclude the survey with a discussion of key unresolved questions in ICRL that can effectively foster a bridge between theoretical understanding and practical industrial applications. The papers referenced in this survey can be found at https://github.com/Jasonxu1225/Awesome-Constraint-Inference-in-RL.

强化学习约束推断综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。