构建城市交通通用模型,实现多视角安全推理。
Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

- 基于开放视觉问答数据集,统一微观驾驶与宏观交通分析
- 在11.6K样本上实现多图像风险分析与目标定位的领先性能
- 适合智能交通、自动驾驶与城市安全研究者参考
城市交通系统面临日益严峻的安全挑战,亟需可扩展的智能能力以支撑新型智慧交通基础设施。尽管基础模型与大规模多模态数据集的发展提升了智能交通系统(ITS)的感知与推理能力,但现有研究仍主要集中于微观自动驾驶(AD),对城市尺度交通分析关注不足。特别是面向安全的开放式视觉问答(VQA)及其对应的基础模型,在异构路侧摄像头观测下的推理仍属空白。为此,我们提出道路交通运输数据集(LTD),一个大规模开源的视觉语言数据集,支持城市交通环境中的开放式推理。LTD包含11.6K条高质量的视觉问答对,来自不同道路形态、交通参与者、光照条件和恶劣天气的异构路侧摄像头。数据集整合了三项互补任务:细粒度多对象定位、多图像相机选择与多图像风险分析,要求在低相关性视图间联合推理以识别危险物体、诱因及高危路段方向。为保障标注质量,采用多模型生成结合交叉验证与人工闭环修正。基于LTD,我们进一步提出UniVLT——一种通过课程式知识迁移训练的交通基础模型,将微观自动驾驶推理与宏观交通分析统一于单一架构中。在LTD及多个自动驾驶基准上的实验表明,UniVLT在跨领域开放式推理任务中达到当前最优表现,同时揭示了现有基础模型在复杂多视角交通场景中的局限性。
原文摘要 · Abstract (English)
Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation models and large-scale multimodal datasets have strengthened perception and reasoning in intelligent transportation systems (ITS), existing research remains largely centered on microscopic autonomous driving (AD), with limited attention to city-scale traffic analysis. In particular, open-ended safety-oriented visual question answering (VQA) and corresponding foundation models for reasoning over heterogeneous roadside camera observations remain underexplored. To address this gap, we introduce the Land Transportation Dataset (LTD), a large-scale open-source vision-language dataset for open-ended reasoning in urban traffic environments. LTD contains 11.6K high-quality VQA pairs collected from heterogeneous roadside cameras, spanning diverse road geometries, traffic participants, illumination conditions, and adverse weather. The dataset integrates three complementary tasks: fine-grained multi-object grounding, multi-image camera selection, and multi-image risk analysis, requiring joint reasoning over minimally correlated views to infer hazardous objects, contributing factors, and risky road directions. To ensure annotation fidelity, we combine multi-model vision-language generation with cross-validation and human-in-the-loop refinement. Building upon LTD, we further propose UniVLT, a transportation foundation model trained via curriculum-based knowledge transfer to unify microscopic AD reasoning and macroscopic traffic analysis within a single architecture. Extensive experiments on LTD and multiple AD benchmarks demonstrate that UniVLT achieves SOTA performance on open-ended reasoning tasks across diverse domains, while exposing limitations of existing foundation models in complex multi-view traffic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。