arXiv:2503.08162cs.ROcs.CL2025-03被引 12

用快慢双系统提升自动驾驶安全,慢系统遇不确定时自动介入

FASIONAD++ : Integrating High-Level Instruction and Information Bottleneck in FAt-Slow fusION Systems for Enhanced Safety in Autonomous Driving with Adaptive Feedback

  • 快系统实时生成轨迹,慢系统在不确定时启动进行深度推理
  • 碰撞率降低28.1%,轨迹误差平均减少6.7%
  • 适合关注自动驾驶安全性与决策可解释性的研究者

确保自动驾驶系统安全、舒适且高效规划至关重要。尽管基于大规模数据训练的端到端模型在常规场景表现良好,但在复杂低频事件中仍显不足。近期大语言模型(LLMs)和视觉语言模型(VLMs)的发展提升了推理能力,但存在计算效率低下问题。受“快思考-慢思考”双过程认知模型启发,我们提出FASIONAD——一种融合快速端到端规划器与VLM推理模块的新型双系统框架。快速系统通过端到端学习实现常见场景下的实时轨迹生成;慢系统则通过不确定性估计触发,执行上下文分析与复杂场景求解。架构引入三大创新:(1) 基于实时不确定性评估的动态切换机制;(2) 带有高层计划反馈的信息瓶颈,优化慢系统引导能力;(3) 双向知识交互,视觉提示增强慢系统推理,其反馈反向优化快系统决策。为强化VLM推理,我们设计问答机制并结合奖励指令训练策略。开环实验显示,FASIONAD实现平均L2轨迹误差降低6.7%,碰撞率下降28.1%。

原文摘要 · Abstract (English)

Ensuring safe, comfortable, and efficient planning is crucial for autonomous driving systems. While end-to-end models trained on large datasets perform well in standard driving scenarios, they struggle with complex low-frequency events. Recent Large Language Models (LLMs) and Vision Language Models (VLMs) advancements offer enhanced reasoning but suffer from computational inefficiency. Inspired by the dual-process cognitive model "Thinking, Fast and Slow", we propose $\textbf{FASIONAD}$ -- a novel dual-system framework that synergizes a fast end-to-end planner with a VLM-based reasoning module. The fast system leverages end-to-end learning to achieve real-time trajectory generation in common scenarios, while the slow system activates through uncertainty estimation to perform contextual analysis and complex scenario resolution. Our architecture introduces three key innovations: (1) A dynamic switching mechanism enabling slow system intervention based on real-time uncertainty assessment; (2) An information bottleneck with high-level plan feedback that optimizes the slow system's guidance capability; (3) A bidirectional knowledge exchange where visual prompts enhance the slow system's reasoning while its feedback refines the fast planner's decision-making. To strengthen VLM reasoning, we develop a question-answering mechanism coupled with reward-instruct training strategy. In open-loop experiments, FASIONAD achieves a $6.7\%$ reduction in average $L2$ trajectory error and $28.1\%$ lower collision rate.

自动驾驶双系统VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。