arXiv:2603.22315cs.LGcs.AI2026-03

用决策变压器优化应急车辆通行,无需实时交互即可灵活控制响应速度。

Emergency Preemption Without Online Exploration: A Decision Transformer Approach

  • 将路径优化建模为离线返回条件序列生成,避免在线环境交互。
  • 在4×4网格上使应急车平均通行时间减少37.7%,仅需1.2次停靠。
  • 支持调度端通过目标返回值调节优先级,适合城市交通系统部署。

应急车辆(EV)响应时间是影响生存率的关键因素,但现有信号优先策略仍为被动响应且不可控。本文基于决策变压器(DT)提出一种返回条件框架,将应急走廊优化建模为离线、返回条件的序列建模任务。该方法(1)在策略学习中完全消除在线环境交互,(2)通过单个目标返回标量实现调度层面的紧急程度控制,(3)通过图注意力机制的多智能体决策变压器(MADT)扩展至多智能体场景,实现空间协调。在LightSim仿真器上,DT在4×4网格中将应急车平均通行时间从142.3秒降至88.6秒(减少37.7%),同时实现最低民用延误(11.3秒/车)和最少停靠次数(1.2次),优于需要环境交互的在线强化学习基线。MADT在8×8网格进一步提升,通过图注意力协调实现45.2%的通行时间降低。返回条件带来平滑调度接口:目标返回值从100调至-400时,应急车通行时间由72.4秒升至138.2秒,民用延误则由16.8秒/车降至5.4秒/车,无需重新训练。进一步的约束型DT引入显式民用干扰预算作为第二控制变量。

原文摘要 · Abstract (English)

Emergency vehicle (EV) response time is a critical determinant of survival outcomes, yet deployed signal preemption strategies remain reactive and uncontrollable. We propose a return-conditioned framework for emergency corridor optimization based on the Decision Transformer (DT). By casting corridor optimization as offline, return-conditioned sequence modeling, our approach (1) eliminates online environment interaction during policy learning, (2) enables dispatch-level urgency control through a single target-return scalar, and (3) extends to multi-agent settings via a Multi-Agent Decision Transformer (MADT) with graph attention for spatial coordination. On the LightSim simulator, DT reduces average EV travel time by 37.7% relative to fixed-timing preemption on a 4x4 grid (88.6 s vs. 142.3 s), achieving the lowest civilian delay (11.3 s/veh) and fewest EV stops (1.2) among all methods, including online RL baselines that require environment interaction. MADT further improves on larger grids, overtaking DT with 45.2% reduction on 8x8 via graph-attention coordination. Return conditioning produces a smooth dispatch interface: varying the target return from 100 to -400 trades EV travel time (72.4-138.2 s) against civilian delay (16.8-5.4 s/veh), requiring no retraining. A Constrained DT extension adds explicit civilian disruption budgets as a second control knob.

决策变换应急调度交通优化多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。