用单一大模型统一处理交通多模态数据,提升决策效率与精度。
Multimodal LLM for Intelligent Transportation Systems
- 构建三维框架,用一个数据驱动的大模型融合时序、图像与视频数据。
- 在多个交通数据集上平均准确率达91.33%,时序数据最高达92.7%。
- 适合希望简化模型部署的交通智能系统研发者参考。
在交通系统智能化演进中,引入大语言模型(LLM)为智能决策提供了新可能。本文提出一种三维框架,整合应用场景、机器学习方法与硬件设备,特别强调LLM的作用。不同于使用多种算法,该框架采用单一数据驱动的LLM架构,可同时处理时间序列、图像与视频数据。研究在Oxford Radar RobotCar、D-Behavior(D-Set)、nuScenes(Motional)、Comma2k19等传感器数据集上验证,旨在简化数据处理流程,降低多模型部署复杂度,提升交通系统的效率与准确性。实验基于AMD RTX 3060 GPU与Intel i9-12900处理器进行,结果表明,该框架在所有数据集上平均准确率达91.33%,其中时序数据最高达92.7%,充分展现其对序列信息的处理能力,适用于运动规划与预测性维护等任务。研究证实了LLM在交通领域处理多模态数据的普适性与有效性,为真实场景应用提供可行路径。
原文摘要 · Abstract (English)
In the evolving landscape of transportation systems, integrating Large Language Models (LLMs) offers a promising frontier for advancing intelligent decision-making across various applications. This paper introduces a novel 3-dimensional framework that encapsulates the intersection of applications, machine learning methodologies, and hardware devices, particularly emphasizing the role of LLMs. Instead of using multiple machine learning algorithms, our framework uses a single, data-centric LLM architecture that can analyze time series, images, and videos. We explore how LLMs can enhance data interpretation and decision-making in transportation. We apply this LLM framework to different sensor datasets, including time-series data and visual data from sources like Oxford Radar RobotCar, D-Behavior (D-Set), nuScenes by Motional, and Comma2k19. The goal is to streamline data processing workflows, reduce the complexity of deploying multiple models, and make intelligent transportation systems more efficient and accurate. The study was conducted using state-of-the-art hardware, leveraging the computational power of AMD RTX 3060 GPUs and Intel i9-12900 processors. The experimental results demonstrate that our framework achieves an average accuracy of 91.33\% across these datasets, with the highest accuracy observed in time-series data (92.7\%), showcasing the model's proficiency in handling sequential information essential for tasks such as motion planning and predictive maintenance. Through our exploration, we demonstrate the versatility and efficacy of LLMs in handling multimodal data within the transportation sector, ultimately providing insights into their application in real-world scenarios. Our findings align with the broader conference themes, highlighting the transformative potential of LLMs in advancing transportation technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。