轻量级多模态模型让自动驾驶系统更易更新与评估
LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving
- 基于视觉语言模型构建统一框架,支持动态更新与快速验证
- 在nuScenes数据集上测试12个代理,发现复杂度提升不等于性能提升
- 适合关注自动驾驶系统可迭代设计的研究者与开发者
视觉语言模型(VLMs)在端到端自动驾驶中展现出巨大潜力,但现有领域缺乏能支持动态模型更新、快速验证、公平比较和直观性能评估的实用平台。为此,我们提出LightEMMA——一种轻量级端到端多模态自动驾驶模型。LightEMMA提供无需定制化的统一VLM驱动框架,可便捷集成不断演进的商用及开源模型。我们基于多种VLM构建了12个自动驾驶代理,在具有挑战性的nuScenes预测任务上进行评估,全面分析计算指标并提供关键洞察。示例显示,尽管VLM具备强场景理解能力,其在实际自动驾驶任务中的表现仍存疑。此外,模型复杂度增加与推理时长延长,并未必然带来性能提升,凸显出进一步改进与任务专用设计的必要性。代码已公开于https://github.com/michigan-traffic-lab/LightEMMA。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated significant potential for end-to-end autonomous driving. However, the field still lacks a practical platform that enables dynamic model updates, rapid validation, fair comparison, and intuitive performance assessment. To that end, we introduce LightEMMA, a Lightweight End-to-End Multimodal Model for Autonomous driving. LightEMMA provides a unified, VLM-based autonomous driving framework without ad hoc customizations, enabling easy integration with evolving state-of-the-art commercial and open-source models. We construct twelve autonomous driving agents using various VLMs and evaluate their performance on the challenging nuScenes prediction task, comprehensively assessing computational metrics and providing critical insights. Illustrative examples show that, although VLMs exhibit strong scenario interpretation capabilities, their practical performance in autonomous driving tasks remains a concern. Additionally, increased model complexity and extended reasoning do not necessarily lead to better performance, emphasizing the need for further improvements and task-specific designs. The code is available at https://github.com/michigan-traffic-lab/LightEMMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。