一个能处理多模态数据的通用自动驾驶大模型
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
- 统一模型处理图像与多视角视频,支持感知、预测、规划任务
- 在6个公开基准上达领先性能,零样本迁移3个新数据集仍优秀
- 适合追求端到端通用解决方案的自动驾驶研究者
大型多模态模型(LMM)通过融合大语言模型,在自动驾驶中展现出卓越的理解与解析能力。然而,现有数据驱动方法通常局限于单一数据集和特定任务,忽视了整体能力与泛化性。为此,我们提出RoboTron-Drive,一个通用的大规模多模态模型,可处理图像与多视角视频等多样化输入,执行感知、预测与规划等广泛任务。模型首先通过课程预训练学习多种视觉信号与基础理解能力;随后对多个自动驾驶数据集进行增强与标准化,完成微调,形成全栈式自动驾驶大模型。为评估其通用性与泛化能力,我们在六个公开基准上进行测试,并在三个未见数据集上进行零样本迁移,结果表明RoboTron-Drive在所有任务上均达到当前最佳性能。我们希望该模型能为真实世界自动驾驶提供有效解决方案。项目页面及代码:https://github.com/zhijian11/RoboTron-Drive。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current data-driven AD approaches tend to concentrate on a single dataset and specific tasks, neglecting their overall capabilities and ability to generalize. To bridge these gaps, we propose RoboTron-Drive, a general large multimodal model designed to process diverse data inputs, such as images and multi-view videos, while performing a broad spectrum of AD tasks, including perception, prediction, and planning. Initially, the model undergoes curriculum pre-training to process varied visual signals and perform basic visual comprehension and perception tasks. Subsequently, we augment and standardize various AD datasets to finetune the model, resulting in an all-in-one LMM for autonomous driving. To assess the general capabilities and generalization ability, we conduct evaluations on six public benchmarks and undertake zero-shot transfer on three unseen datasets, where RoboTron-Drive achieves state-of-the-art performance across all tasks. We hope RoboTron-Drive as a promising solution for AD in the real world. Project page with code: https://github.com/zhijian11/RoboTron-Drive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。