发现最优策略有简单几何结构,直接学习决策区域更高效。
Low-Complexity Policy Tessellations in Structured Markov Decision Processes

- 基于边界直接学习策略区域,避免复杂值函数近似。
- 策略误差更低,值差距更小,收敛更快且更稳定。
- 适合需要高效决策的库存与排队控制场景。
我们研究结构化马尔可夫决策过程中的最优策略几何特性。尽管近似动态规划和强化学习通常需近似高维值函数,我们发现最优策略会诱导出更简单的决策剖分。提出基于边界的策略近似方法,直接学习策略区域。通过策略损失分解,揭示性能下降与动作边际的关系,解释为何误差集中在无差异边界附近。库存控制和队列准入实验表明,该方法相比强化学习基线,具有更低的策略误差、更小的值差距、更快的误差衰减速度以及更高的稳定性。
原文摘要 · Abstract (English)
We study optimal-policy geometry in structured Markov decision processes. While approximate dynamic programming and reinforcement learning typically approximate high-dimensional value functions, we show that optimal policies induce simpler decision tessellations. We propose boundary-based policy approximations that learn policy regions directly. A policy-loss decomposition links performance degradation to action margins and explains why errors concentrate near indifference boundaries. Inventory control and queue admission experiments show lower policy error, smaller value gaps, faster error decay, and stability than reinforcement learning baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。