用智能体框架解决快递最后一公里订单决策冲突问题。
ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery

- 构建决策点模型,显式建模骑手与订单的时空关系和权衡。
- 通过结构化报告发现排序分歧,提升决策可解释性。
- 在四城数据上比现有方法平均提升9.2%,适合配送系统优化者。
最后一公里配送需应对动态到达的订单与骑手之间的复杂时空关联。现有学习方法虽能预测骑手服务序列,但对下一订单选择缺乏解释。通过语言描述当前配送状态,大模型可显式推理空间、时间及行为线索。然而,直接使用大模型作为预测器易受任务呈现方式影响,导致决策不可靠。为此,我们提出ORBITER——一个用于最后一公里配送中下一订单决策的智能体订单仲裁系统。ORBITER以决策点形式建模骑手服务过程,每个决策点包含骑手的时空状态及可见订单,并暴露局部权衡以供建模与验证。固定提案者对候选订单排序,结构化报告揭示排序差异。大模型利用任务特定工具收集主要候选方案的证据,独立评判者则基于证据核查最终决策。我们在四个城市的实测数据上进行了广泛评估,结果显示ORBITER相较现有最先进基线平均提升9.2%,证明其有效性。
原文摘要 · Abstract (English)
Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recent learning-based methods model spatiotemporal dependencies among orders to predict courier service sequences, but leave next-order decision making unexplained. Describing the current delivery state in language allows LLMs to reason explicitly about the spatial, temporal, and behavioral cues behind an individual decision. As direct predictors, however, LLMs remain sensitive to task presentation and often produce unreliable decisions. To address these challenges, we introduce ORBITER, an agentic Order Arbiter for next-order decision-making in last-mile delivery. ORBITER models courier service through decision points, each containing the courier's spatiotemporal state and visible orders and exposing local trade-offs for modeling and verification. Fixed proposers rank the candidates, and a structured report identifies where their rankings disagree. The LLM uses task-specific tools to gather evidence on the leading alternatives, while an independent critic checks the resulting decision against that evidence. We conduct extensive evaluations on data in four cities, where ORBITER outperforms existing state-of-the-art baselines by up to 9.2% on average showing its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。