arXiv:2608.10897cs.LG2026-08

解决多平台配送中信息不全的难题,提升真实场景下的调度效率。

Partially Observable Learning for Multi-Platform Dispatch Optimization

论文配图:Partially Observable Learning for Multi-Platform Dispatch Optimization
图 1 · 摘自论文原文
  • 基于局部观测构建独立智能体,模拟真实隐私限制下的决策。
  • 引入注意力机制聚合骑手信息,应对信息不完整与异构性挑战。
  • 适合研究多平台协同调度或实际物流系统优化的开发者。

即时配送平台已成为城市物流的关键环节,越来越多依赖众包骑手完成高度动态的订单。现实中,骑手并不专属于单一平台,可能同时为多个平台服务,而各平台仅能观测自身订单及骑手互动,受隐私和运营约束导致多平台调度环境具有固有的部分可观测性。现有大多数调度优化方法假设骑手信息完全可观测且必须接受任务,导致在真实多平台环境中性能严重下降。本文提出POLO,一种面向多平台即时配送系统的部分可观测多智能体强化学习框架。POLO将每个平台-网格对建模为独立智能体,仅基于本平台本地观测学习调度策略,符合现实中的隐私与运营约束。为支持在不完整且异构的骑手信息下做出有效决策,POLO设计了一种新颖的基于注意力的策略表示,可选择性聚合骑手间信息。此外,我们提出了反事实奖励重塑机制,缓解跨网格联合动作带来的非平稳性,实现更稳定、可扩展的学习。我们构建了一个高保真仿真器,评估不同平台数量与系统规模下的调度表现。大量实验表明,POLO在平台收益与骑手出行效率上均持续优于强基线,凸显其在真实多平台场景下的鲁棒性与有效性。

原文摘要 · Abstract (English)

Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders. In real-world systems, couriers are not exclusive to a single platform and may concurrently serve multiple platforms, while each platform can only observe its own orders and couriers' interactions due to privacy and operational constraints. This results in a multi-platform dispatch environment with inherent partial observability. However, most existing works on dispatch optimization assume full courier observability and mandatory assignment acceptance, causing substantial performance degradation when deployed in realistic multi-platform settings. In this paper, we propose POLO, a partially observable multi-agent reinforcement learning framework for dispatching optimization in multi-platform instant delivery systems. POLO firstly models each platform-grid pair as an independent agent that learns dispatch policies solely from platform-local observations, aligning the learning process with real-world privacy and operational constraints. To support effective decision-making under incomplete and heterogeneous courier information, POLO introduces a novel attention-based policy representation that selectively aggregates inter-courier information. Moreover, we design a counterfactual reward shaping mechanism to mitigate the non-stationarity induced by joint actions across grids, leading to more stable and scalable learning. We develop a high-fidelity simulator to evaluate dispatch performance under varying numbers of platforms and system scales. Extensive experiments demonstrate that POLO consistently outperforms strong baselines in terms of platform revenue and courier travel efficiency, highlighting its robustness and effectiveness in realistic multi-platform settings.

调度优化多智能体强化学习即时配送

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。