arXiv:2502.12756cs.LGmath.OC2025-02中稿 · KDD被引 1

用深度强化学习解决海运集装箱配载的不确定性难题。

Navigating Demand Uncertainty in Container Shipping: Deep Reinforcement Learning for Enabling Adaptive and Feasible Master Stowage Planning

  • 设计编码器-解码器框架融合船况与需求信息指导决策
  • 通过可微投影层在复杂约束下保持解的可行性
  • 可在需求波动时自适应调整,适合物流系统应用

强化学习在解决确定性和随机规划问题上已取得成功,但传统方法难以应对现实世界中依赖状态的复杂约束。本文以海运集装箱主配载规划为实际案例,研究在需求不确定和运营约束下的随机序列决策问题。提出一种基于编码器-解码器结构的深度强化学习框架,整合问题实例、解空间与不确定性信息以引导规划。引入可微投影层实现凸多面体约束的强制满足,并通过雅可比修正补偿投影偏差,获得无偏策略梯度估计。实验表明,该模型能高效生成自适应且可行的解,在分布偏移下具有泛化能力,可扩展至更长规划周期,性能优于当前最先进的约束强化学习与随机规划方法。所提策略支持对不确定性敏感的动态规划,有助于构建韧性与可持续的供应链。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has successfully solved various deterministic and stochastic planning problems. However, conventional RL struggles with complex real-world constraints, particularly when feasibility is explicit and depends on the current state or trajectory. In this work, we address stochastic sequential decision-making with state-dependent constraints through a real-world case study of the master stowage planning problem in container shipping, which aims to optimize revenue and costs under demand uncertainty and operational constraints. We propose a deep RL framework with an encoder-decoder model that integrates problem instance, solution, and uncertainty information to guide planning. We introduce differentiable projection layers that enforce convex polyhedral constraints, while Jacobian corrections offset the projections to yield unbiased policy gradient estimates. Experiments show that our model efficiently finds adaptive, feasible solutions that generalize across distribution shifts and scale to longer planning horizons, outperforming state-of-the-art baselines in constrained RL and stochastic programming. As such, our policies enable adaptive, uncertainty-aware planning that can support resilient and sustainable supply chains.

强化学习航运优化不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。