让智能体不仅学专家平均表现,还学会其风险偏好。
Imitation Learning as Return Distribution Matching
- 用瓦瑟斯坦距离匹配专家回报分布,实现风险敏感模仿学习。
- 提出两种算法,利用动态信息使样本效率显著提升。
- 非马尔可夫策略优于传统方法,适合高风险决策场景。
我们研究通过模仿学习训练风险敏感的强化学习智能体。与标准模仿学习不同,目标不仅是匹配专家的期望回报(即平均表现),还包括其风险态度(如回报分布的方差等特征)。我们提出一个通用的风险敏感模仿学习框架,目标是通过瓦瑟斯坦距离匹配专家的回报分布。在表格设置下,假设专家奖励已知,我们证明了马尔可夫策略在此任务中的表达能力有限,因此引入一种高效且足够丰富的非马尔可夫策略子类。基于该子类,我们设计出两种可证明高效的算法:当转移模型未知时使用RS-BC,已知时使用RS-KT。结果表明,利用动态信息的RS-KT相比RS-BC显著降低样本复杂度。此外,在专家奖励未知情况下,通过设计基于预言机的RS-KT变体,进一步验证了回报分布匹配的样本效率。最后,数值实验支持理论分析,凸显了非马尔可夫策略相对于标准高效模仿学习算法的优势。
原文摘要 · Abstract (English)
We study the problem of training a risk-sensitive reinforcement learning (RL) agent through imitation learning (IL). Unlike standard IL, our goal is not only to train an agent that matches the expert's expected return (i.e., its average performance) but also its risk attitude (i.e., other features of the return distribution, such as variance). We propose a general formulation of the risk-sensitive IL problem in which the objective is to match the expert's return distribution in Wasserstein distance. We focus on the tabular setting and assume the expert's reward is known. After demonstrating the limited expressivity of Markovian policies for this task, we introduce an efficient and sufficiently expressive subclass of non-Markovian policies tailored to it. Building on this subclass, we develop two provably efficient algorithms, RS-BC and RS-KT, for solving the problem when the transition model is unknown and known, respectively. We show that RS-KT achieves substantially lower sample complexity than RS-BC by exploiting dynamics information. We further demonstrate the sample efficiency of return distribution matching in the setting where the expert's reward is unknown by designing an oracle-based variant of RS-KT. Finally, we complement our theoretical analysis of RS-KT and RS-BC with numerical simulations, highlighting both their sample efficiency and the advantages of non-Markovian policies over standard sample-efficient IL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。