arXiv:2508.09183quant-phcs.AI2025-08被引 2

用量子增强强化学习,优化实时配送的路径规划。

Quantum-Efficient Reinforcement Learning Solutions for Last-Mile On-Demand Delivery

  • 设计量子电路与强化学习结合的新框架,处理大规模配送问题。
  • 在真实时间窗约束下,相比传统方法显著降低旅行时间。
  • 适合关注量子计算与物流优化交叉应用的研究者。

量子计算为求解NP难组合优化问题提供了有前景的替代方案。当面对大规模优化时,经典方法难以应对。本文研究利用量子计算解决具有时间窗约束的大规模带容量取送货问题(CPDPTW)。为此,设计了一种融合参数化量子电路(PQC)的强化学习(RL)框架,以最小化实际最后一公里按需配送中的行程时间。提出一种新型问题专用编码量子电路,包含纠缠与变分层结构。通过数值实验对比了近端策略优化(PPO)和量子奇异值变换(QSVT),结果表明所提方法在解的规模和训练复杂度上均具优势,且能有效融入现实约束条件。

原文摘要 · Abstract (English)

Quantum computation has demonstrated a promising alternative to solving the NP-hard combinatorial problems. Specifically, when it comes to optimization, classical approaches become intractable to account for large-scale solutions. Specifically, we investigate quantum computing to solve the large-scale Capacitated Pickup and Delivery Problem with Time Windows (CPDPTW). In this regard, a Reinforcement Learning (RL) framework augmented with a Parametrized Quantum Circuit (PQC) is designed to minimize the travel time in a realistic last-mile on-demand delivery. A novel problem-specific encoding quantum circuit with an entangling and variational layer is proposed. Moreover, Proximal Policy Optimization (PPO) and Quantum Singular Value Transformation (QSVT) are designed for comparison through numerical experiments, highlighting the superiority of the proposed method in terms of the scale of the solution and training complexity while incorporating the real-world constraints.

量子计算强化学习路径优化物流配送

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。