arXiv:2603.28963cs.ROcs.AI2026-03

用原始激光雷达数据训练交通仿真模型,提升真实感与泛化能力。

AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models

  • 基于自监督世界模型,从原始点云学习交通行为
  • 在部分遮挡场景下表现更优,轨迹生成更贴近真实
  • 适合自动驾驶仿真、多智能体系统研究者

真实交通代理的仿真对验证自动驾驶系统至关重要。现有数据驱动仿真器依赖感知模块输出的3D边界框和折线等高层抽象,这些信息损失了直接影响代理行为的感官上下文,限制了仿真的分布真实性。为此,我们提出AutoWorld,一种基于自监督世界模型的交通仿真框架,该模型在LiDAR占用数据上训练,将代理行为锚定于原始传感器观测。给定世界模型采样结果,AutoWorld构建粗到细的预测场景上下文作为多智能体运动生成模型的输入。此外,我们设计了一种运动感知的隐空间监督目标,丰富了对场景动态的隐表示。为更好地利用该隐空间进行推理,AutoWorld采用级联行列式点过程框架,实现世界模型与运动模型间的多样性感知采样。在Waymo Sim Agents Challenge (WOSAC) 上的实验表明,AutoWorld达到有竞争力的性能,在部分观测场景中提升尤为显著。进一步显示,通过原始LiDAR建模的AutoWorld在数据量增加时扩展性优于仅使用轨迹或仅依赖LiDAR的基线方法。消融实验验证了各组件的有效性。

原文摘要 · Abstract (English)

Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agent behavior from higher-level abstractions such as 3D bounding boxes and polylines, inferred by upstream perception pipelines. These lossy abstractions discard sensory context that directly shapes agent behavior, limiting the distributional realism that simulation aims to reproduce. To address this limitation, we propose AutoWorld, a traffic simulation framework that grounds agent behavior in raw sensor observations through a self-supervised world model trained on LiDAR occupancy data. Given world model samples, AutoWorld constructs a coarse-to-fine predictive scene context as input to a multi-agent motion generation model. Furthermore, we designed a motion-aware latent supervision objective that enriches AutoWorld's latent representation of scene dynamics. To better exploit this latent space during inference, AutoWorld employs a cascaded Determinantal Point Process framework to guide diversity-aware sampling across both the world model and motion model. Experiments on the Waymo Sim Agents Challenge (WOSAC) demonstrate that AutoWorld achieves competitive performance, with larger gains in partially-observed scenarios where trajectory abstractions are most limited. We further show that grounding simulation in raw LiDAR through AutoWorld scales better with additional data than trajectory-only and LiDAR-conditioning baselines. Ablations confirm the contribution of each component.

交通仿真自监督多智能体激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。