arXiv:2505.19053cs.LGmath.OC2025-05NeurIPS被引 10

将组合优化层嵌入强化学习,提升复杂决策问题的效率与稳定性。

Structured Reinforcement Learning for Combinatorial Decision-Making

  • 在策略网络中引入组合优化层,利用Fenchel-Young损失实现端到端训练。
  • 在动态任务上性能优于基线最多提升92%,且收敛更快、更稳定。
  • 适合处理路由、调度等具有结构化动作空间的复杂决策问题。

强化学习(RL)正越来越多地应用于涉及复杂且结构化决策的实际问题,如路径规划、调度和产品组合设计。这些场景对标准RL算法构成挑战,因其在组合动作空间下难以扩展、泛化并利用结构。本文提出结构化强化学习(SRL),一种新型的演员-评论家范式,将组合优化层嵌入演员神经网络。通过Fenchel-Young损失实现演员的端到端学习,并提供SRL在矩量多面体对偶空间中作为原始-对偶算法的几何解释。在六个包含外生与内生不确定性的环境中,SRL在静态任务上达到或超过无结构RL与模仿学习的性能,在动态任务上相比基线最高提升92%,同时具备更好的稳定性与更快的收敛速度。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed.

强化学习组合优化决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。