开源多模态模型让自动驾驶更高效可靠
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
- 基于多模态大模型与思维链推理,实现端到端自动驾驶
- 在复杂场景下表现优异,优于基线模型
- 完全开源,适合研究者和开发者快速迭代
自多模态大语言模型(MLLMs)问世以来,其在自动驾驶等现实应用中展现出显著影响力。它们能处理复杂的视觉信息并推理复杂驾驶情境,推动了端到端自动驾驶系统的新范式。然而,现有端到端模型开发进展缓慢,因微调方法需大量计算资源、大规模数据集和资金支持。受推理计算最新进展启发,我们提出OpenEMMA——一个基于MLLMs的开源端到端自动驾驶框架。通过引入思维链(Chain-of-Thought)推理机制,OpenEMMA在多种多模态模型上均实现显著性能提升。此外,该框架在多样挑战性驾驶场景中表现出有效性、泛化性和鲁棒性,提供了一种更高效、更可靠的自动驾驶方案。代码已开源:https://github.com/taco-group/OpenEMMA。
原文摘要 · Abstract (English)
Since the advent of Multimodal Large Language Models (MLLMs), they have made a significant impact across a wide range of real-world applications, particularly in Autonomous Driving (AD). Their ability to process complex visual data and reason about intricate driving scenarios has paved the way for a new paradigm in end-to-end AD systems. However, the progress of developing end-to-end models for AD has been slow, as existing fine-tuning methods demand substantial resources, including extensive computational power, large-scale datasets, and significant funding. Drawing inspiration from recent advancements in inference computing, we propose OpenEMMA, an open-source end-to-end framework based on MLLMs. By incorporating the Chain-of-Thought reasoning process, OpenEMMA achieves significant improvements compared to the baseline when leveraging a diverse range of MLLMs. Furthermore, OpenEMMA demonstrates effectiveness, generalizability, and robustness across a variety of challenging driving scenarios, offering a more efficient and effective approach to autonomous driving. We release all the codes in https://github.com/taco-group/OpenEMMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。