arXiv:2606.18820cs.LGcs.AI2026-06

提出新型马尔可夫决策模型,解决信息变多而可选动作减少的决策难题。

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

论文配图:Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets
图 1 · 摘自论文原文
  • 基于信息与动作的不对称演化构建新模型
  • 实验显示复杂场景下学习效率提升显著
  • 适合大规模动态决策系统研究者参考

序列决策问题常呈现信息与决策灵活性的非对称演化:随着决策周期推进,智能体接收更丰富的信息,但可行动作因操作截止、承诺或资源限制而逐渐消失。标准MDP通常将此结构简化为阶段依赖的状态描述和动作掩码,从而掩盖了决定哪些决策紧迫、哪些可延后的信息-动作不对称性。我们提出成熟马尔可夫决策过程(MMDP),围绕这一不对称性构建。通过一个随时间消逝的动作优先原则,揭示了必须在下一阶段前解决的动作。受此结构启发,我们设计了一种结构感知强化学习框架,包含阶段感知策略、随时间消逝动作抽象以及带蒸馏的搜索增强学习。在可控多供应商补货问题、复杂度递增的简化现金管理环境及生产规模仿真器上的实验表明,显式建模该不对称性能提升学习效率,并在问题规模增大时价值愈发显著。

原文摘要 · Abstract (English)

Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints. Standard MDP formulations typically flatten this structure into stage-dependent state descriptions and action masks, thereby obscuring the nested information--action asymmetry that determines which decisions are urgent and which can be deferred. We introduce Maturing Markov Decision Processes (MMDPs), a formulation built around this information--action asymmetry. We characterize one of its key consequences through an expiring-action priority principle, which identifies the actions that must be resolved before the next stage. Motivated by this structure, we develop a structure-aware reinforcement learning framework with stage-aware policy design, expiring-action abstraction, and search-augmented learning with distillation. Experiments on a controlled multi-supplier replenishment problem, simplified cash-management environments of increasing complexity, and a production-scale simulator show that explicitly modeling this asymmetry improves learning efficiency and becomes increasingly valuable as decision problems scale.

强化学习决策优化动态规划信息不对称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。