arXiv:2605.02461stat.MLcs.LG2026-05被引 1

用目标导向强化学习优化中段物流路径,提升运输效率。

Middle-mile logistics through the lens of goal-conditioned reinforcement learning

  • 将中段物流建模为多目标条件马尔可夫决策过程
  • 结合图神经网络与无模型强化学习,提取环境特征图
  • 适合物流算法研究者和智能调度系统开发者

中段物流指通过有限容量卡车连接的枢纽网络对包裹进行路由的问题。本文将其重新表述为一个多目标条件马尔可夫决策过程(MDP)。方法结合图神经网络与无模型强化学习,从环境状态中提取小型特征图,以高效建模复杂物流网络中的动态调度问题。

原文摘要 · Abstract (English)

Middle-mile logistics describes the problem of routing parcels through a network of hubs linked by trucks with finite capacity. We rephrase this as a multi-object goal-conditioned MDP. Our method combines graph neural networks with model-free RL, extracting small feature graphs from the environment state.

物流优化强化学习图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。