arXiv:2603.11093cs.AIcs.RO2026-03综述被引 1

自动驾驶正从感知转向认知,需构建可解释的推理核心应对复杂场景挑战。

A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms

  • 提出认知层级框架,按认知与交互复杂度拆解驾驶任务
  • 识别出七项关键推理挑战,如响应与推理的权衡问题
  • 强调未来需发展可验证的神经符号架构以连接符号推理与物理控制

自动驾驶的发展正从感知局限转向更根本的瓶颈:缺乏鲁棒且泛化的推理能力。当前系统虽能处理结构化环境,但在长尾场景和复杂社交互动中仍表现不佳,难以实现人类般的判断力。大语言模型(LLMs)与多模态模型(MLLMs)的兴起为引入强大的认知引擎提供了契机,推动系统从模式匹配迈向真正理解。然而,指导这一融合的系统性框架仍严重缺失。为此,本文全面综述该新兴领域,主张将推理提升为系统的认知核心。首先提出一种新的认知层级框架,依据认知与交互复杂度分解驾驶任务;在此基础上,归纳并系统化七项核心推理挑战,包括响应-推理权衡、社会博弈推理等。进一步从系统架构与评估方法双重视角分析前沿进展,揭示向整体化、可解释的“玻璃盒”智能体演进的趋势。最后指出:基于大模型的推理具有高延迟、强思考特性,与车辆控制毫秒级安全需求存在根本矛盾。未来重点应是弥合符号与物理之间的鸿沟,发展可验证的神经符号架构、不确定环境下的稳健推理及支持隐式社会协商的可扩展模型。

原文摘要 · Abstract (English)

The development of high-level autonomous driving (AD) is shifting from perception-centric limitations to a more fundamental bottleneck, namely, a deficit in robust and generalizable reasoning. Although current AD systems manage structured environments, they consistently falter in long-tail scenarios and complex social interactions that require human-like judgment. Meanwhile, the advent of large language and multimodal models (LLMs and MLLMs) presents a transformative opportunity to integrate a powerful cognitive engine into AD systems, moving beyond pattern matching toward genuine comprehension. However, a systematic framework to guide this integration is critically lacking. To bridge this gap, we provide a comprehensive review of this emerging field and argue that reasoning should be elevated from a modular component to the system's cognitive core. Specifically, we first propose a novel Cognitive Hierarchy to decompose the monolithic driving task according to its cognitive and interactive complexity. Building on this, we further derive and systematize seven core reasoning challenges, such as the responsiveness-reasoning trade-off and social-game reasoning. Furthermore, we conduct a dual-perspective review of the state-of-the-art, analyzing both system-centric approaches to architecting intelligent agents and evaluation-centric practices for their validation. Our analysis reveals a clear trend toward holistic and interpretable "glass-box" agents. In conclusion, we identify a fundamental and unresolved tension between the high-latency, deliberative nature of LLM-based reasoning and the millisecond-scale, safety-critical demands of vehicle control. For future work, a primary objective is to bridge the symbolic-to-physical gap by developing verifiable neuro-symbolic architectures, robust reasoning under uncertainty, and scalable models for implicit social negotiation.

自动驾驶推理系统大模型认知架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。