通过模仿学习生成满足约束的安全策略,理论严谨且训练稳定。
Learning safe, constrained policies via imitation learning: Connection to Probabilistic Inference and a Naive Algorithm
- 基于概率推断框架,将约束与奖励统一于最大熵目标
- 采用对偶梯度下降优化,支持多类型约束下的有效训练
- 可泛化至不同行为模式,适合安全敏感的强化学习场景
本文提出一种模仿学习方法,用于学习符合专家轨迹所展示约束的最大熵策略。该方法利用演示策略与学习策略间KL散度的性能界结果,其目标函数通过强化学习的概率推断框架严格推导,将奖励最大化与约束遵守统一在熵最大化设置中。所提算法采用对偶梯度下降优化学习目标,实现高效稳定的训练。实验表明,该方法可在包含多种类型约束、不同行为模态及具备泛化能力的设置下,成功学习出满足约束的有效策略。
原文摘要 · Abstract (English)
This article introduces an imitation learning method for learning maximum entropy policies that comply with constraints demonstrated by expert trajectories executing a task. The formulation of the method takes advantage of results connecting performance to bounds for the KL-divergence between demonstrated and learned policies, and its objective is rigorously justified through a connection to a probabilistic inference framework for reinforcement learning, incorporating the reinforcement learning objective and the objective to abide by constraints in an entropy maximization setting. The proposed algorithm optimizes the learning objective with dual gradient descent, supporting effective and stable training. Experiments show that the proposed method can learn effective policy models for constraints-abiding behaviour, in settings with multiple constraints of different types, accommodating different modalities of demonstrated behaviour, and with abilities to generalize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。