arXiv:2501.14394cs.LG2025-01

用强化学习即时分配退货商品,大幅缩短仓储时间。

Reinforcement Learning for Efficient Returns Management

  • 将退货分配转为在线多重背包问题,实时决策
  • 存储时间减少96%,收益仅比离线方法低3%
  • 适合追求高效仓储的零售企业

在零售仓库中,退货商品通常需暂存于中间仓,等待后续配送至门店。存放时间越长,管理成本与效率损失越高,因需持续占用存储空间。为缩短平均存储时间,本文提出一种新方案:商品到仓后立即决策再分配,仅允许有限数量商品同时暂存。将该问题转化为在线多重背包问题,设计了一种新型强化学习算法,以最大化整体预期收益为目标,将商品(物品)分配至门店(背包)。在模拟数据上的实验表明,相比传统的离线决策方式,本方法性能差距仅为3%,但商品平均存储时间却降低了96%。

原文摘要 · Abstract (English)

In retail warehouses, returned products are typically placed in an intermediate storage until a decision regarding further shipment to stores is made. The longer products are held in storage, the higher the inefficiency and costs of the returns management process, since enough storage area has to be provided and maintained while the products are not placed for sale. To reduce the average product storage time, we consider an alternative solution where reallocation decisions for products can be made instantly upon their arrival in the warehouse allowing only a limited number of products to still be stored simultaneously. We transfer the problem to an online multiple knapsack problem and propose a novel reinforcement learning approach to pack the items (products) into the knapsacks (stores) such that the overall value (expected revenue) is maximized. Empirical evaluations on simulated data demonstrate that, compared to the usual offline decision procedure, our approach comes with a performance gap of only 3% while significantly reducing the average storage time of a product by 96%.

强化学习仓储优化退货管理在线决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。