用模型预测控制提升强化学习的效率与安全性
Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents
- 结合模型预测控制与强化学习,利用系统先验知识优化决策
- 在保证安全性和可解释性的同时,显著减少样本需求
- 适合需要高可靠性的自动驾驶、机器人控制场景
在不确定性环境下实现最优决策是现代自主系统发展的关键。尽管无模型强化学习(RL)通过系统交互直接提升性能,且对系统先验知识要求低,但其依赖深度神经网络导致样本效率低、学习过程不安全、决策不可解释。为此,本文提出以模型为基础的智能体作为替代方案,利用系统动力学、代价和约束的可调模型,实现安全策略学习。这些模型可融入先验知识,指导、约束并解释智能体决策,而模型偏差可通过无模型强化学习修正。文中阐述了基于模型的智能体(如模型预测控制)的优势与挑战,详细介绍了贝叶斯优化、策略搜索强化学习和离线策略等主要学习方法及其各自优势。尽管无模型强化学习已成熟,其与基于模型方法的协同潜力尚未充分探索,本文旨在揭示二者结合在实现高效、安全、可解释决策智能体中的前景。
原文摘要 · Abstract (English)
Training sophisticated agents for optimal decision-making under uncertainty has been key to the rapid development of modern autonomous systems across fields. Notably, model-free reinforcement learning (RL) has enabled decision-making agents to improve their performance directly through system interactions, with minimal prior knowledge about the system. Yet, model-free RL has generally relied on agents equipped with deep neural network function approximators, appealing to the networks' expressivity to capture the agent's policy and value function for complex systems. However, neural networks amplify the issues of sample inefficiency, unsafe learning, and limited interpretability in model-free RL. To this end, this work introduces model-based agents as a compelling alternative for control policy approximation, leveraging adaptable models of system dynamics, cost, and constraints for safe policy learning. These models can encode prior system knowledge to inform, constrain, and aid in explaining the agent's decisions, while deficiencies due to model mismatch can be remedied with model-free RL. We outline the benefits and challenges of learning model-based agents -- exemplified by model predictive control -- and detail the primary learning approaches: Bayesian optimization, policy search RL, and offline strategies, along with their respective strengths. While model-free RL has long been established, its interplay with model-based agents remains largely unexplored, motivating our perspective on their combined potentials for sample-efficient learning of safe and interpretable decision-making agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。