用瑞尼散度统一风险敏感与最大熵强化学习
Risk-sensitive control as inference with Rényi divergence
- 以瑞尼散度构建变分推断框架,扩展经典控制即推断方法
- 风险敏感策略可通过软贝尔曼方程求解,保持与最优后验一致
- 适用于需权衡风险与收益的决策场景,如金融、机器人控制
本文提出风险敏感控制即推断(RCaI),通过瑞尼散度变分推断扩展经典控制即推断(CaI)。RCaI被证明等价于对数概率正则化的风险敏感控制,是最大熵(MaxEnt)控制的推广。我们进一步证明风险敏感最优策略可通过求解软贝尔曼方程获得,揭示了RCaI、MaxEnt控制、CaI最优后验及线性可解控制之间的多重等价关系。基于此,我们推导出风险敏感强化学习方法:策略梯度与软演员-评论家算法。当风险敏感参数趋近于零时,可恢复风险中性形式的CaI与RL,表明RCaI是一个统一框架。此外,我们还提出了基于瑞尼熵正则化的另一种最大熵控制推广。尽管推导路径不同,两种扩展的最优策略具有相同结构。
原文摘要 · Abstract (English)
This paper introduces the risk-sensitive control as inference (RCaI) that extends CaI by using Rényi divergence variational inference. RCaI is shown to be equivalent to log-probability regularized risk-sensitive control, which is an extension of the maximum entropy (MaxEnt) control. We also prove that the risk-sensitive optimal policy can be obtained by solving a soft Bellman equation, which reveals several equivalences between RCaI, MaxEnt control, the optimal posterior for CaI, and linearly-solvable control. Moreover, based on RCaI, we derive the risk-sensitive reinforcement learning (RL) methods: the policy gradient and the soft actor-critic. As the risk-sensitivity parameter vanishes, we recover the risk-neutral CaI and RL, which means that RCaI is a unifying framework. Furthermore, we give another risk-sensitive generalization of the MaxEnt control using Rényi entropy regularization. We show that in both of our extensions, the optimal policies have the same structure even though the derivations are very different.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。