用LLM动态构建环境知识库,提升自动驾驶感知与决策能力
SenseRAG: Constructing Environmental Knowledge Bases with Proactive Querying for LLM-Based Autonomous Driving
- 通过主动检索增强生成,将多模态数据融合为LLM可读知识库
- 在真实V2X数据上实现感知与预测性能显著提升
- 适合关注智能驾驶系统推理能力的开发者与研究者
本研究针对自动驾驶中情境感知能力不足的问题,利用大语言模型(LLMs)的上下文推理能力,提出一种基于主动检索增强生成(Proactive RAG)的框架。该框架不依赖传统固定标签的感知系统,而是将实时多模态传感器数据整合为统一、可被LLM理解的知识库,支持动态环境理解与响应。为克服LLM固有的延迟与模态限制,引入链式思维提示机制,确保快速且富含上下文的理解。基于真实世界车路协同(Vehicle-to-everything, V2X)数据集的实验表明,该方法在感知与预测任务上均有显著性能提升,展现出提升下一代自动驾驶系统安全性、适应性与决策能力的巨大潜力。
原文摘要 · Abstract (English)
This study addresses the critical need for enhanced situational awareness in autonomous driving (AD) by leveraging the contextual reasoning capabilities of large language models (LLMs). Unlike traditional perception systems that rely on rigid, label-based annotations, it integrates real-time, multimodal sensor data into a unified, LLMs-readable knowledge base, enabling LLMs to dynamically understand and respond to complex driving environments. To overcome the inherent latency and modality limitations of LLMs, a proactive Retrieval-Augmented Generation (RAG) is designed for AD, combined with a chain-of-thought prompting mechanism, ensuring rapid and context-rich understanding. Experimental results using real-world Vehicle-to-everything (V2X) datasets demonstrate significant improvements in perception and prediction performance, highlighting the potential of this framework to enhance safety, adaptability, and decision-making in next-generation AD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。