将订单履约拆解为准备与交付两阶段,提升动态环境下的配送效率。
A Policy Decomposition Framework for Dynamic Order Fulfillment Operations
- 通过策略分解将交付阶段独立优化,准备阶段仅作状态约束。
- 在真实数据集上优于独立或联合优化的基线方法,显著降低延迟。
- 适合需要高动态响应的电商与定制制造供应链场景。
现代供应链涵盖从电商配送网络到按单生产等多种运营环境。其高效运作依赖于订单准备与下游交付两个高度耦合阶段的协同。尽管传统上二者被分开管理,但实际履约系统需在随机到达的订单流下满足严格的交付要求。为此,我们提出动态订单履约问题(DOFP),统一此前分散研究的物流挑战。将DOFP建模为马尔可夫决策过程,状态与决策空间分为准备与交付子空间,由同步约束连接。现有方法多在短视滚动时域内联合优化两阶段,而本框架将下游交付策略独立优化,将准备视为状态级约束过滤器。我们提出基于价值函数近似的分解驱动框架(DDF-VFA),采用新型策略级分解:将搜索划分为交付阶段主问题与准备阶段兼容性子问题,通过反馈循环迭代优化。该框架结合大规模邻域搜索与神经网络价值函数近似,用于估计剩余成本。在两种基于真实数据集的变体上进行数值验证,结果表明DDF-VFA始终优于独立或联合优化的基准方法。此外,该框架可自然扩展以处理批量准备或多阶段准备等现实复杂性。
原文摘要 · Abstract (English)
Modern supply chains span diverse operational environments, ranging from e-commerce distribution networks to customized production-to-order manufacturing lines. Across these settings, operational efficiency depends on coordinating two highly interdependent stages: order preparation and downstream delivery. Although these stages are traditionally managed in isolation, real-world fulfillment systems must satisfy stringent delivery expectations under dynamic stochastic order arrivals. To bridge this gap, we introduce the Dynamic Order Fulfillment Problem (DOFP), a new problem class unifying logistical challenges previously studied separately. We model DOFP as a Markov decision process whose state and decision spaces are partitioned into preparation and delivery sub-spaces, linked by synchronization constraints. While recent approaches attempt to optimize both fulfillment stages simultaneously over myopic rolling horizons, our framework isolates and optimizes the downstream delivery policy, treating preparation strictly as a state-level constraint filter. To solve this, we develop the Decomposition-Driven Framework with Value Function Approximation (DDF-VFA), which utilizes a novel policy-level decomposition. This design partitions the search into a delivery-stage master problem and a preparation-stage compatibility subproblem, iteratively refined via feedback loops. DDF-VFA executes this strategy by combining a large-neighborhood search over partial delivery decisions with a neural-network value function approximation for the cost-to-go. Numerical illustrations on two example variants using real-world datasets show that DDF-VFA consistently outperforms benchmarks that optimize the two stages independently or jointly without decomposition. Finally, the framework naturally scales to accommodate additional real-world complexities such as batched or multi-stage preparation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。