arXiv:2506.12232cs.CVcs.CL2025-06被引 1

用多模态大模型实现自动驾驶零样本场景理解

Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles

  • 在零样本条件下测试四种多模态大模型的场景理解能力
  • GPT-4o表现最佳,但小模型差距不大
  • 集成方法效果不一,需更优融合策略

场景理解对自动驾驶的下游任务至关重要,如促进人机交互和提升车辆决策的可解释性。本文评估了四种多模态大语言模型(MLLMs),包括较小模型,在零样本、上下文学习设置下的场景理解能力。同时,探索通过多数投票集成方法是否能提升性能。实验表明,最大的模型GPT-4o表现最优,但与小型模型的差距相对有限,提示可通过改进上下文学习、检索增强生成(RAG)或微调等技术进一步优化小模型。集成方法结果参差不齐:部分场景属性的F1分数有所提升,另一些则下降。这表明需发展更复杂的集成策略以在所有场景属性上实现一致提升。本研究凸显了利用MLLM进行场景理解的潜力,并为自动驾驶应用中的模型优化提供了洞见。

原文摘要 · Abstract (English)

Scene understanding is critical for various downstream tasks in autonomous driving, including facilitating driver-agent communication and enhancing human-centered explainability of autonomous vehicle (AV) decisions. This paper evaluates the capability of four multimodal large language models (MLLMs), including relatively small models, to understand scenes in a zero-shot, in-context learning setting. Additionally, we explore whether combining these models using an ensemble approach with majority voting can enhance scene understanding performance. Our experiments demonstrate that GPT-4o, the largest model, outperforms the others in scene understanding. However, the performance gap between GPT-4o and the smaller models is relatively modest, suggesting that advanced techniques such as improved in-context learning, retrieval-augmented generation (RAG), or fine-tuning could further optimize the smaller models' performance. We also observe mixed results with the ensemble approach: while some scene attributes show improvement in performance metrics such as F1-score, others experience a decline. These findings highlight the need for more sophisticated ensemble techniques to achieve consistent gains across all scene attributes. This study underscores the potential of leveraging MLLMs for scene understanding and provides insights into optimizing their performance for autonomous driving applications.

多模态大模型自动驾驶零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。