arXiv:2411.10603cs.ROcs.SY2024-11被引 2

用GPT-4o增强自动驾驶在恶劣天气下的决策能力。

A Novel MLLM-based Approach for Autonomous Driving in Different Weather Conditions

  • 基于MLLM的自动驾驶代理,融合GPT-4o与CARLA仿真环境。
  • 在雨雾等恶劣条件下仍保持高安全与效率表现。
  • 适合研究多模态大模型在智能驾驶中的应用者。

自动驾驶技术有望通过提升安全性、效率与舒适性,彻底改变日常交通。在恶劣环境条件下,其面临重大挑战,亟需鲁棒且自适应的解决方案。本文提出一种基于多模态大语言模型(MLLM)的自动驾驶代理 MLLM-AD-4o,利用 GPT-4o 在 LimSim++ 框架中实现与 CARLA 驾驶模拟器的闭环交互。该研究评估了代理在雨雾、低能见度及复杂交通场景下的感知、决策与控制能力。结果表明,该代理在极端环境下仍能维持高水平的安全性与效率,凸显 GPT-4o 增强自动驾驶系统(ADS)的潜力。同时,实验对比了仅前视相机、前后相机组合以及融合 LiDAR 的不同感知配置,为未来集成 MLLM 与自动驾驶框架提供了关键见解。

原文摘要 · Abstract (English)

Autonomous driving (AD) technology promises to revolutionize daily transportation by making it safer, more efficient, and more comfortable. Their role in reducing traffic accidents and improving mobility will be vital to the future of intelligent transportation systems. Autonomous driving in harsh environmental conditions presents significant challenges that demand robust and adaptive solutions and require more investigation. In this context, we present in this paper a comprehensive performance analysis of an autonomous driving agent leveraging the capabilities of a Multi-modal Large Language Model (MLLM) using GPT-4o within the LimSim++ framework that offers close loop interaction with the CARLA driving simulator. We call it MLLM-AD-4o. Our study evaluates the agent's decision-making, perception, and control under adverse conditions, including bad weather, poor visibility, and complex traffic scenarios. Our results demonstrate the AD agent's ability to maintain high levels of safety and efficiency, even in challenging environments, underscoring the potential of GPT-4o to enhance autonomous driving systems (ADS) in any environment condition. Moreover, we evaluate the performance of MLLM-AD-4o when different perception entities are used including either front cameras only, front and rear cameras, and when combined with LiDAR. The results of this work provide valuable insights into integrating MLLMs with AD frameworks, paving the way for future advancements in this field.

自动驾驶多模态GPT-4o仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。