提升自动驾驶在视野受限区的推理能力,融合感知与常识知识
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
- 设计指令引导交互模块,降低输入长度并连接多模态信息
- 构建200万条问答数据集,支持基于常识的多步风险推理
- 适用于提升复杂交通场景下对弱势道路使用者的安全判断
具备广泛世界知识的多模态大语言模型(MLLMs)正在重塑自动驾驶中的推理任务,尤其在可感知区域表现突出。然而,在感知受限区域(如动态或静态遮挡区域),MLLMs难以有效融合感知能力与世界知识进行推理。这些区域可能隐藏关键安全信息,尤其是对弱势道路使用者。本文提出一种框架,旨在通过增强感知能力与世界知识的融合,提升自动驾驶在感知受限条件下的表现。具体而言,设计了一个即插即用的指令引导交互模块,有效弥合模态差异并显著减少输入序列长度,可适应多视角视频输入。为进一步融合世界知识与驾驶任务,构建并精炼了大规模多模态数据集,包含200万条自然语言问答对和170万条定位任务数据。为评估模型对世界知识的利用能力,引入一个对象级风险评估数据集,包含20万条需多步推理才能解答的问答对。大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perceptionlimited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model's utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。