用大模型+多传感器融合,让自动驾驶更懂路况、决策更可靠。
DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
- 分角色协作:摄像头、激光雷达等传感器由专用代理分析,大模型统筹决策。
- 在多个数据集上超越基线,关键场景理解与动作规划更准确。
- 适合研究自动驾驶智能体、多模态融合的开发者和工程师参考。
我们提出DriveAgent,一种新型多智能体自动驾驶框架,结合大语言模型(LLM)推理与多模态传感器融合,提升对交通环境的理解与决策能力。该框架将相机、激光雷达、GPS和惯性测量单元(IMU)等多种传感器数据,通过四个模块化智能体协同处理:(i) 描述性分析代理根据筛选的时间戳识别关键事件;(ii) 激光雷达与视觉代理联合分析车辆状态与运动;(iii) 环境推理与因果分析代理解释上下文变化及其成因;(iv) 时效感知决策生成代理整合信息并建议及时操作。这种模块化设计使大模型能高效协调感知与推理任务,输出连贯且可解释的驾驶判断。在多个挑战性自动驾驶数据集上的实验表明,DriveAgent在多项指标上优于基线方法,验证了其在提升自动驾驶系统鲁棒性与可靠性方面的有效性。
原文摘要 · Abstract (English)
We introduce DriveAgent, a novel multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion to enhance situational understanding and decision-making. DriveAgent uniquely integrates diverse sensor modalities-including camera, LiDAR, GPS, and IMU-with LLM-driven analytical processes structured across specialized agents. The framework operates through a modular agent-based pipeline comprising four principal modules: (i) a descriptive analysis agent identifying critical sensor data events based on filtered timestamps, (ii) dedicated vehicle-level analysis conducted by LiDAR and vision agents that collaboratively assess vehicle conditions and movements, (iii) environmental reasoning and causal analysis agents explaining contextual changes and their underlying mechanisms, and (iv) an urgency-aware decision-generation agent prioritizing insights and proposing timely maneuvers. This modular design empowers the LLM to effectively coordinate specialized perception and reasoning agents, delivering cohesive, interpretable insights into complex autonomous driving scenarios. Extensive experiments on challenging autonomous driving datasets demonstrate that DriveAgent is achieving superior performance on multiple metrics against baseline methods. These results validate the efficacy of the proposed LLM-driven multi-agent sensor fusion framework, underscoring its potential to substantially enhance the robustness and reliability of autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。