优化线性系统的互信息,同时学习最优策略和先验分布。
Mutual Information Optimal Control of Discrete-Time Linear Systems
- 联合优化控制策略与先验分布,突破传统固定先验的限制。
- 在高斯分布假设下,推导出策略与先验的最优解形式。
- 提出交替优化算法,实验验证其有效性和收敛性。
本文提出了离散时间线性系统的互信息最优控制问题(MIOCP),可视为最大熵最优控制问题(MEOCP)的扩展。与MEOCP中先验固定为均匀分布不同,MIOCP同时优化策略与先验。在策略与先验均属于高斯分布类的前提下,我们分别推导出固定先验时的最优策略和固定策略时的最优先验。基于这些解析结果,提出一种交替最小化算法求解MIOCP。通过数值实验,分析了所提算法的工作机制与性能表现。
原文摘要 · Abstract (English)
In this paper, we formulate a mutual information optimal control problem (MIOCP) for discrete-time linear systems. This problem can be regarded as an extension of a maximum entropy optimal control problem (MEOCP). Differently from the MEOCP where the prior is fixed to the uniform distribution, the MIOCP optimizes the policy and prior simultaneously. As analytical results, under the policy and prior classes consisting of Gaussian distributions, we derive the optimal policy and prior of the MIOCP with the prior and policy fixed, respectively. Using the results, we propose an alternating minimization algorithm for the MIOCP. Through numerical experiments, we discuss how our proposed algorithm works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。