arXiv:2412.02025cs.ROcs.AI2024-12中稿 · presentation at IC…被引 23

用统一提示框架让大模型像人一样逐步思考,提升自动驾驶决策能力。

PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving

  • 设计基于感知、知识、推理、决策的零样本思维链提示
  • GPT-4在多个任务中表现优异,无需额外训练
  • 适用于多种大模型,适合快速验证自动驾驶方案

随着多模态大语言模型(MLLMs)在自动驾驶中的应用日益广泛,端到端模型高昂的开发成本和复杂性限制了其普及。为此,本文提出一种零样本思维链提示方法PKRD-CoT,基于自动驾驶的四大核心能力:感知、知识、推理与决策。该方法使大模型能在无先验经验的情况下,模拟人类逐步思考过程,有效应对动态驾驶环境,提升实时决策能力。实验表明,GPT-4配合PKRD-CoT在多项自动驾驶任务中表现出色;基准测试还验证了其对Claude、LLaVA1.6、Qwen-VL-Plus等模型的通用性。本研究为GPT-4及其他主流MLLMs提供了可复用的统一提示框架,并通过全面对比验证了其在自动驾驶场景下的有效性。

原文摘要 · Abstract (English)

There is growing interest in leveraging the capabilities of robust Multi-Modal Large Language Models (MLLMs) directly within autonomous driving contexts. However, the high costs and complexity of designing and training end-to-end autonomous driving models make them challenging for many enterprises and research entities. To address this, our study explores a seamless integration of MLLMs into autonomous driving systems by proposing a Zero-Shot Chain-of-Thought (Zero-Shot-CoT) prompt design named PKRD-CoT. PKRD-CoT is based on the four fundamental capabilities of autonomous driving: perception, knowledge, reasoning, and decision-making. This makes it particularly suitable for understanding and responding to dynamic driving environments by mimicking human thought processes step by step, thus enhancing decision-making in real-time scenarios. Our design enables MLLMs to tackle problems without prior experience, thereby increasing their utility within unstructured autonomous driving environments. In experiments, we demonstrate the exceptional performance of GPT-4.0 with PKRD-CoT across autonomous driving tasks, highlighting its effectiveness in autonomous driving scenarios. Additionally, our benchmark analysis reveals the promising viability of PKRD-CoT for other MLLMs, such as Claude, LLava1.6, and Qwen-VL-Plus. Overall, this study contributes a novel and unified prompt-design framework for GPT-4.0 and other MLLMs in autonomous driving, while also rigorously evaluating the efficacy of these widely recognized MLLMs in the autonomous driving domain through comprehensive comparisons.

自动驾驶大模型提示工程思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。