arXiv:2409.05061cs.AI2024-09被引 2

动态调控快递柜可用性,提升配送效率与客户满意度

Dynamic Demand Management for Parcel Lockers

  • 根据包裹优先级动态决定是否开放柜子接收服务
  • 相比基准方法提升13.7%的可服务请求量
  • 适合关注物流优化与智能调度的研究者

为实现更可持续、低成本的末端配送,快递柜已在包裹配送中占据重要地位。为充分发挥其潜力并保障客户满意度,对有限柜体容量的有效管理至关重要。由于未来配送请求和取件时间具有不确定性,管理难度增加。为此,本文提出动态控制每个新客户是否可选择使用快递柜,以最大化加权服务请求数。同时考虑不同隔间尺寸,需对已计划配送的包裹进行分配决策。将问题建模为无限时域序贯决策问题,发现精确方法因维度灾难难以求解。因此,我们构建一个融合顺序决策分析与强化学习的技术框架,包括成本函数近似、离线训练的参数化价值函数近似及截断在线滚动策略。创新性地结合多种技术,有效处理两类决策间的强关联。方法论上,采用改进的经验回放机制增强价值函数结构。计算实验表明,本方法相较贪心基准提升13.7%,较行业启发式策略提升12.6%。

原文摘要 · Abstract (English)

In pursuit of a more sustainable and cost-efficient last mile, parcel lockers have gained a firm foothold in the parcel delivery landscape. To fully exploit their potential and simultaneously ensure customer satisfaction, successful management of the locker's limited capacity is crucial. This is challenging as future delivery requests and pickup times are stochastic from the provider's perspective. In response, we propose to dynamically control whether the locker is presented as an available delivery option to each incoming customer with the goal of maximizing the number of served requests weighted by their priority. Additionally, we take different compartment sizes into account, which entails a second type of decision as parcels scheduled for delivery must be allocated. We formalize the problem as an infinite-horizon sequential decision problem and find that exact methods are intractable due to the curses of dimensionality. In light of this, we develop a solution framework that orchestrates multiple algorithmic techniques rooted in Sequential Decision Analytics and Reinforcement Learning, namely cost function approximation and an offline trained parametric value function approximation together with a truncated online rollout. Our innovative approach to combine these techniques enables us to address the strong interrelations between the two decision types. As a general methodological contribution, we enhance the training of our value function approximation with a modified version of experience replay that enforces structure in the value function. Our computational study shows that our method outperforms a myopic benchmark by 13.7% and an industry-inspired policy by 12.6%.

物流优化强化学习决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。