arXiv:2512.01518cs.LG2025-12被引 2

用端到端强化学习解决带不确定性的多目标车辆路径问题。

End-to-end Deep Reinforcement Learning for Stochastic Multi-objective Optimization in C-VRPTW

  • 基于注意力机制与多轨迹的深度强化学习模型
  • 在合理时间内生成高质量帕累托解集,优于三个基线方法
  • 适合需要兼顾效率与合规的物流调度场景

本文研究基于学习的路由方法,解决具有随机性和多目标特征的车辆路径问题。实际场景中决策者需应对运营环境的不确定性及不同利益相关方的目标冲突。我们聚焦旅行时间不确定性,同时优化总行程时间和路线完工时间,以兼顾运营效率与工时法规。由于旅行时间矩阵维度高,无法直接输入模型,因此难以处理不确定性,尤其在多目标情形下更具挑战性。为此,我们提出一种同时处理随机性和多目标的端到端深度学习模型,并引入场景聚类优化训练过程以减少训练时间。实验结果表明,该模型在可接受的运行时间内生成了高质量的帕累托前沿,显著优于三种基线方法。

原文摘要 · Abstract (English)

In this work, we consider learning-based applications in routing to solve a Vehicle Routing variant characterized by stochasticity and multiple objectives. Such problems are representative of practical settings where decision-makers have to deal with uncertainty in the operational environment as well as multiple conflicting objectives due to different stakeholders. We specifically consider travel time uncertainty. We also consider two objectives, total travel time and route makespan, that jointly target operational efficiency and labor regulations on shift length, although different objectives could be incorporated. Learning-based methods offer earnest computational advantages as they can repeatedly solve problems with limited interference from the decision-maker. We specifically focus on end-to-end deep learning models that leverage the attention mechanism and multiple solution trajectories. These models have seen several successful applications in routing problems. However, since travel times are not a direct input to these models due to the large dimensions of the travel time matrix, accounting for uncertainty is a challenge, especially in the presence of multiple objectives. In turn, we propose a model that simultaneously addresses stochasticity and multi-objectivity and provide a refined training mechanism for this model through scenario clustering to reduce training time. Our results show that our model is capable of constructing a Pareto Front of good quality within acceptable run times compared to three baselines.

强化学习车辆路径多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。