构建可控交通场景生成与评估平台,推动时空交通推理研究。
OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

- 基于12个真实路口重建3D环境,支持可控生成多视角交通数据。
- 生成800万同步VQA样本,测试集经3000人验证,揭示模型严重短板。
- 适合交通视觉、多模态大模型、自动驾驶系统研究者使用。
交通场景理解需超越物体识别,涵盖车道拓扑、多视角几何、时序演化与信号相位语义。现有交通多模态基准大多聚焦被动视觉识别或孤立视频理解,难以在可控条件下评估结构感知的交通推理能力。本文提出OmniTraffic——一个可控制的生成流水线与基准,基于12个真实路口重建的可编辑3D交通环境,并结合两国监控视频,支持可控与自然条件下的评估。其定义三级任务体系:场景感知、多视角与时序推理、决策支持。利用结构化交通元数据,生成同步多视角VQA样本,覆盖车辆状态、车道功能、视图-俯视图对应、时序动态及信号相位分析,共生成800万样本,构建3000人验证的测试集。对十一款前沿多模态大模型(MLLM)的评估显示显著的人机差距,尤其在拓扑依赖和时空推理任务中表现最差。在模拟的OmniTraffic数据上微调轻量级MLLM后,其在真实交通场景中的表现得到提升,证明了仿真生成监督对交通特定多模态推理的价值。OmniTraffic不仅是一个固定数据集,更提供可配置的扩展流水线,支持自定义路口、相机视角、交通需求、信号相位、视觉条件及罕见事件。
原文摘要 · Abstract (English)
Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. However, existing traffic-oriented multimodal benchmarks largely emphasize passive visual recognition or isolated video understanding, offering limited support for evaluating structure-aware traffic reasoning under controlled conditions. We introduce OmniTraffic, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built around 12 real-world intersections reconstructed into editable 3D traffic environments and complemented by surveillance footage from two countries, OmniTraffic supports both controlled and natural-condition evaluation. It defines a three-level task hierarchy spanning scene perception, multi-view and temporal reasoning, and decision support. Using structured traffic metadata, OmniTraffic generates synchronized multi-view VQA samples covering vehicle states, lane functions, view--BEV correspondence, temporal dynamics, and signal-phase analysis, resulting in 8M VQA samples and a 3K human-verified test set. Evaluation of eleven frontier MLLMs reveals a large human--model gap, with the most pronounced failures in topology-grounded and spatio-temporal reasoning tasks. Fine-tuning a lightweight MLLM on simulated OmniTraffic data further improves performance on real-world traffic scenes, demonstrating the value of simulation-generated supervision for traffic-specific multimodal reasoning. Beyond a fixed dataset, OmniTraffic provides an extensible pipeline with configurable intersections, camera views, traffic demands, signal phases, visual conditions, and rare events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。