揭示线性系统中互信息控制的随机性与温度参数的关系
On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
- 通过理论推导建立温度参数与策略随机性的关联
- 给出策略变随机或确定的温度阈值条件
- 适用于控制理论研究者和强化学习算法设计者
近年来,互信息最优控制作为最大熵最优控制的扩展被提出。两种方法均引入正则项使策略具有随机性,但温度参数(即正则项系数)与策略随机性的理论关系在互信息最优控制中尚未明确。本文研究离散时间线性系统的互信息最优控制问题(MIOCP),在拓展先前研究结果的基础上,证明了MIOCP最优策略的存在性,并推导出最优策略为随机或确定的温度参数条件。同时,还推导了交替优化算法所得策略在何种温度下变为随机或确定。数值实验验证了理论结果的有效性。
原文摘要 · Abstract (English)
In recent years, mutual information optimal control has been proposed as an extension of maximum entropy optimal control. Both approaches introduce regularization terms to render the policy stochastic, and it is important to theoretically clarify the relationship between the temperature parameter (i.e., the coefficient of the regularization term) and the stochasticity of the policy. Unlike in maximum entropy optimal control, this relationship remains unexplored in mutual information optimal control. In this paper, we investigate this relationship for a mutual information optimal control problem (MIOCP) of discrete-time linear systems. After extending the result of a previous study of the MIOCP, we establish the existence of an optimal policy of the MIOCP, and then derive the respective conditions on the temperature parameter under which the optimal policy becomes stochastic and deterministic. Furthermore, we also derive the respective conditions on the temperature parameter under which the policy obtained by an alternating optimization algorithm becomes stochastic and deterministic. The validity of the theoretical results is demonstrated through numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。