arXiv:2604.10664cs.AI2026-04被引 1

动态调整优先级的实时车辆调度优化方法

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

论文配图:Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching
图 1 · 摘自论文原文
  • 基于深度强化学习构建可动态输入偏好向量的统一模型
  • 在码头车辆调度任务中性能优于两种主流MOO方法
  • 适合需要实时调整目标权重的复杂动态决策场景

多目标优化(MOO)因其在真实场景中支持以人为本的决策而被广泛研究。随着市场动态加剧,对动态MOO的需求迅速增长,要求实时调整不同目标的优先级。然而,现有研究或聚焦于非现实的确定性MOO问题,或处理非序列化的动态决策问题,难以应对实际复杂性。为此,本文提出偏好敏捷多目标优化(PAMOO),支持用户动态调整并实时交互设定偏好。通过在深度强化学习(DRL)框架内设计新型统一模型,显式接收用户动态偏好向量作为输入,并引入校准函数确保偏好向量与DRL决策策略之间的高质量对齐。在集装箱码头的复杂实时车辆调度任务上进行的大量实验表明,与两种最流行的MOO方法相比,PAMOO展现出更优的性能和泛化能力。本方法首次解决具有挑战性的动态序列多目标决策问题。

原文摘要 · Abstract (English)

Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life applications. Recently, demand for dynamic MOO is fast-emerging due to tough market dynamics that require real-time re-adjustments of priorities for different objectives. However, most existing studies focus either on deterministic MOO problems which are not practical, or non-sequential dynamic MOO decision problems that cannot deal with some real-life complexities. To address these challenges, a preference-agile multi-objective optimization (PAMOO) is proposed in this paper to permit users to dynamically adjust and interactively assign the preferences on the fly. To achieve this, a novel uniform model within a deep reinforcement learning (DRL) framework is proposed that can take as inputs users' dynamic preference vectors explicitly. Additionally, a calibration function is fitted to ensure high quality alignment between the preference vector inputs and the output DRL decision policy. Extensive experiments on challenging real-life vehicle dispatching problems at a container terminal showed that PAMOO obtains superior performance and generalization ability when compared with two most popular MOO methods. Our method presents the first dynamic MOO method for challenging \rev{dynamic sequential MOO decision problems

多目标优化强化学习实时调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。