用Transformer实现千种商品实时补货决策,比传统方法快400万倍。
OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items
- 设计可置换等变的Transformer架构,适应海量商品组合决策
- 在1024种商品场景下性能超越基线方法,且规模越大优势越明显
- 适合需要实时大规模供应链优化的工业场景
现代供应链需在相关随机需求、异构提前期和共享固定订货成本下协调上千种异质商品的补货,观测空间维度超10⁴。在此规模下,滚动时域随机混合整数规划(MILP)求解过于缓慢,而标准强化学习面临高维动作空间中的信用分配难题。本文提出OR-Transformer,一种基于深度强化学习的联合补货框架,采用物品置换等变的Transformer结构,并通过库存动态进行路径梯度训练。在最多1,024种库存商品的问题规模下,OR-Transformer随规模增长持续优于学习型与滚动时域MILP基线。其在线决策时间相比MILP求解器降低超过400万倍,实现了供应链运营中大规模深度强化学习的实时化。
原文摘要 · Abstract (English)
Modern supply chain operations can require coordinating replenishment across thousands of heterogeneous items under correlated stochastic demand, heterogeneous lead times, and shared fixed ordering costs, yielding observation spaces exceeding $10^4$ dimensions. At this scale, rolling-horizon stochastic mixed-integer linear programs (MILPs) become prohibitively slow, while standard reinforcement learning (RL) methods face increasingly challenging credit assignment in high-dimensional action spaces. We introduce OR-Transformer, a deep reinforcement learning framework for joint replenishment under stochastic demand, with an item-permutation-equivariant Transformer architecture and pathwise-gradient training through the inventory dynamics. Across problem sizes up to 1,024 inventory items, OR-Transformer increasingly outperforms learning-based and rolling-horizon MILP baselines as scale grows. It also reduces online decision-making time by over 4 million times relative to MILP solvers, enabling real-time, large-scale deep RL in supply chain operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。