arXiv:2606.21726cs.AI2026-06

为随机系统设计稳定状态分布的控制策略,提升系统可预测性。

Entropy Objectives in Markov Decision Processes

  • 基于熵约束构建马尔可夫决策过程的策略合成方法
  • 提出可验证且相对完备的算法,解决非线性目标难题
  • 适合关注系统稳定性与可控性的自动化设计研究者

我们研究如何合成控制策略,以确保随机系统的状态分布具有集中性。本文在马尔可夫决策过程(MDP)框架下形式化了基于熵的目标策略合成问题。首先证明,即使放宽条件,该问题仍属于计算复杂性难题。随后,提出一种可靠且(在条件下)相对完备的方法,用于验证与合成满足熵目标的策略。核心挑战在于目标函数的非线性特性,我们的方法通过结合凸对偶与不变式合成的思想予以应对。此外,还探讨了记忆与随机性在实现熵目标中的作用。最后,我们在若干典型基准上实现了该方法并进行了实证评估。

原文摘要 · Abstract (English)

We consider the problem of synthesizing control policies that enforce a concentration property on the state distributions of a stochastic system. We present a formalization of this problem in terms of synthesizing strategies for maintaining an entropy-based objective in Markov Decision Processes (MDPs). We first show that even relaxed versions of this problem are complexity-theoretically hard. We then present a sound and (conditionally) relatively complete method to verify and synthesize strategies for such entropy objectives. The main challenge is the non-linear nature of such objectives, and our approach addresses this by exploiting and combining ideas from convex duality and invariant synthesis. We also investigate the role of memory and randomization in ensuring entropy objectives. Finally, we implement our ideas to evaluate our approach empirically on a few illustrative benchmarks.

强化学习控制策略熵优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。