arXiv:2604.17651cs.CVcs.RO2026-04被引 1

让路边传感器生成能预测交通的智能模型,突破自动驾驶视角局限。

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

  • 以路边固定传感器为视角,融合多模态数据构建时空互补的世界模型。
  • 利用长期观测数据捕捉罕见事故场景,实现对复杂交通行为的精准预测。
  • 适合交通管理、智能网联车研发人员,推动车路协同系统进化。

世界模型作为模拟环境演化过程的生成式AI系统,正重塑自动驾驶技术。然而当前所有方法均基于车辆自身视角,忽视了路边基础设施的独特优势。本文提出基础设施中心的世界模型(I-WM),充分利用路边传感器具备的鸟瞰视角、多源感知与持续观测能力。固定传感器擅长时间深度,可积累长期行为分布,包括罕见的安全关键事件;车载传感器则擅长空间广度,覆盖大规模道路网络中的多样场景。本文分三阶段构建I-WM:(I)质量感知的不确定性传播生成场景理解;(II)融合物理规律与多智能体反事实推理的预测动力学;(III)通过潜在空间对齐实现车路协同通信的联合建模。提出双层架构,无需标注的感知作为多模态数据引擎,支持从激光雷达到4D雷达、信号灯相位数据再到事件相机的渐进式部署。建立驾驶世界模型范式分类体系,将I-WM置于LeCun的JEPA、李飞飞的空间智能与VLA架构之间,并引入基础设施视觉语言动作模型(I-VLA),统一路边感知、自然语言指令与交通控制动作。该愿景依托现有多激光雷达管线,识别各阶段开源基础,为构建能理解与预判交通的基础设施提供可行路径。

原文摘要 · Abstract (English)

World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego-vehicle perspective, leaving the infrastructure viewpoint unexplored. We argue that infrastructure-centric world models offer a fundamentally complementary capability: the bird's-eye, multi-sensor, persistent viewpoint that roadside systems uniquely possess. Central to our thesis is a spatio-temporal complementarity: fixed roadside sensors excel at temporal depth, accumulating long-term behavioral distributions including rare safety-critical events, while vehicle-borne sensors excel at spatial breadth, sampling diverse scenes across large road networks. This paper presents a vision for Infrastructure-centric World Models (I-WM) in three phases: (I) generative scene understanding with quality-aware uncertainty propagation, (II) physics-informed predictive dynamics with multi-agent counterfactual reasoning, and (III) collaborative world models for V2X communication via latent space alignment. We propose a dual-layer architecture, annotation-free perception as a multi-modal data engine feeding end-to-end generative world models, with a phased sensor strategy from LiDAR through 4D radar and signal phase data to event cameras. We establish a taxonomy of driving world model paradigms, position I-WM relative to LeCun's JEPA, Li Fei-Fei's spatial intelligence, and VLA architectures, and introduce Infrastructure VLA (I-VLA) as a novel unification of roadside perception, language commands, and traffic control actions. Our vision builds upon existing multi-LiDAR pipelines and identifies open-source foundations for each phase, providing a path toward infrastructure that understands and anticipates traffic.

世界模型车路协同交通预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。