arXiv:2604.05195cs.LG2026-04

用智能提示机制解决多种车辆配送难题,效率高且能泛化到新场景。

Vehicle-as-Prompt: A Unified Deep Reinforcement Learning Framework for Heterogeneous Fleet Vehicle Routing Problem

  • 将车辆作为提示信息输入,统一建模异构车队路径规划
  • 在复杂约束下比现有深度强化学习方法提升显著,推理仅需数秒
  • 零样本迁移能力强,适合实际物流系统快速部署

与传统的同质化路径规划不同,异构车队车辆路径问题(HFVRP)涉及不同的固定成本、可变行驶成本和容量限制,导致解的质量对车辆选择高度敏感。真实物流应用常包含额外复杂约束,显著增加计算难度。然而,现有基于深度强化学习(DRL)的方法多局限于同质场景,在处理HFVRP及其复杂变体时表现不佳。为此,本文研究在复杂约束下的HFVRP,提出统一的DRL框架。引入车辆作为提示(Vehicle-as-Prompt, VaP)机制,将问题建模为单阶段自回归决策过程。在此基础上,提出VaP-CSMV框架,包含跨语义编码器与多视角解码器,有效捕捉车辆异质性与客户节点属性间的复杂映射关系。大量实验表明,VaP-CSMV显著优于现有最优DRL神经求解器,性能媲美传统启发式算法,同时推理时间缩短至数秒。该框架在大规模未见问题变体上展现出强大零样本泛化能力,消融实验证明各组件均具关键贡献。

原文摘要 · Abstract (English)

Unlike traditional homogeneous routing problems, the Heterogeneous Fleet Vehicle Routing Problem (HFVRP) involves heterogeneous fixed costs, variable travel costs, and capacity constraints, rendering solution quality highly sensitive to vehicle selection. Furthermore, real-world logistics applications often impose additional complex constraints, markedly increasing computational complexity. However, most existing Deep Reinforcement Learning (DRL)-based methods are restricted to homogeneous scenarios, leading to suboptimal performance when applied to HFVRP and its complex variants. To bridge this gap, we investigate HFVRP under complex constraints and develop a unified DRL framework capable of solving the problem across various variant settings. We introduce the Vehicle-as-Prompt (VaP) mechanism, which formulates the problem as a single-stage autoregressive decision process. Building on this, we propose VaP-CSMV, a framework featuring a cross-semantic encoder and a multi-view decoder that effectively addresses various problem variants and captures the complex mapping relationships between vehicle heterogeneity and customer node attributes. Extensive experimental results demonstrate that VaP-CSMV significantly outperforms existing state-of-the-art DRL-based neural solvers and achieves competitive solution quality compared to traditional heuristic solvers, while reducing inference time to mere seconds. Furthermore, the framework exhibits strong zero-shot generalization capabilities on large-scale and previously unseen problem variants, while ablation studies validate the vital contribution of each component.

路径优化强化学习物流调度异构车队

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。