用Transformer构建交通行为模型,生成长时间物理一致的轨迹。
Enactor: From Traffic Simulators to Surrogate World Models
- 基于Transformer的代理中心生成模型,融合几何与交互信息
- 4000秒长时仿真中轨迹物理一致性显著提升,训练样本减少
- 适合交通仿真、自动驾驶测试等需要真实交互的场景
交通微观仿真器广泛用于评估道路网络在各种‘假设情景’下的表现。然而,控制参与者行为的模型过于简单,难以捕捉真实的参与者-参与者交互。深度学习方法已将车辆和行人建模为对环境(包括车道、信号灯和邻近参与者)做出响应的‘代理’。尽管能有效学习交互,但这些方法在长时间内难以生成物理一致的轨迹,且未明确处理交通节点这一城市网络关键位置的复杂动态。受世界模型范式启发,我们开发了一种基于Transformer的代理中心生成模型,能够同时捕捉参与者-参与者交互并理解交通节点的几何结构,生成基于学习行为的物理合理轨迹。我们在‘仿真闭环’环境中测试该模型:使用SUMO生成初始条件,再由模型控制参与者动态,仿真运行40000个时间步(4000秒),评估其在长时程下的表现及交通工程相关指标。实验结果表明,所提框架有效捕捉复杂交互,生成长时程物理一致轨迹,且所需训练样本远少于传统代理中心生成方法。模型在交通相关及聚合指标上均优于基线,KL散度指标提升超过10倍。
原文摘要 · Abstract (English)
Traffic microsimulators are widely used to evaluate road network performance under various ``what-if" conditions. However, the behavior models controlling the actions of the actors are overly simplistic and fails to capture realistic actor-actor interactions. Deep learning-based methods have been applied to model vehicles and pedestrians as ``agents" responding to their surrounding ``environment" (including lanes, signals, and neighboring agents). Although effective in learning actor-actor interaction, these approaches fail to generate physically consistent trajectories over long time periods, and they do not explicitly address the complex dynamics that arise at traffic intersections which is a critical location in urban networks. Inspired by the World Model paradigm, we have developed an actor centric generative model using transformer-based architecture that is able to capture the actor-actor interaction, at the same time understanding the geometry to the traffic intersection to generate physically grounded trajectories that are based on learned behavior. Moreover, we test the model in a live ``simulation-in-the-loop" setting, where we generate the initial conditions of the actors using SUMO and then let the model control the dynamics of the actors. We let the simulation run for 40000 timesteps (4000 seconds), testing the performance of the model on long timerange and evaluating the trajectories on traffic engineering related metrics. Experimental results demonstrate that the proposed framework effectively captures complex actor-actor interactions and generates long-horizon, physically consistent trajectories, while requiring significantly fewer training samples than traditional agent-centric generative approaches. Our model is able to outperform the baseline in traffic related as well as aggregate metrics where our model beats the baseline by more than 10x on the KL-Divergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。