arXiv:2605.04475cs.CV2026-05

用统一场景摘要提升自动驾驶感知与推理一致性,减少幻觉。

Information Coordination as a Bridge: A Neuro-Symbolic Architecture for Reliable Autonomous Driving Scene Understanding

  • 构建贝叶斯鸟瞰图为中心的神经符号框架,先融合多传感器数据再推理。
  • 在nuScenes和Waymo上实现98%属性一致率,冗余低于1%。
  • 适合需要高可靠性的自动驾驶系统开发与验证场景。

可靠的自动驾驶需在异构传感器间保持语义一致,并在推理阶段可验证。然而,许多基于大语言模型的系统将语言模型作为后处理模块,迫使它在冗余或冲突的感知输出上推理,易引发幻觉实体和不安全结论。本文提出InfoCoordiBridge,一种以鸟瞰图(BEV)为中心的神经符号架构,在感知与语言推理之间引入显式协调桥接。该架构包含:(i) 统一的多智能体感知层,输出带类型的结构化事实及模态聚焦摘要;(ii) ICA模块,将多源输出对齐并融合为单一场景摘要(SceneSummary);(iii) SSRE模块,基于场景摘要进行可验证推理。在nuScenes和Waymo上的实验表明,ICA在保持竞争性3D检测精度的同时,显著提升融合一致性,冗余降低至1%以下,属性一致率达约98%。在NuScenes-QA与模板对齐的Waymo-QA基准上,SSRE相比代表性视觉语言模型和代理基线,提升了事实准确性,减少了幻觉实体提及。整体而言,通过在提示前将多传感器输出协调为一个冲突感知的场景摘要,InfoCoordiBridge有效防止冗余与跨模态不一致的感知证据传播至高层推理。

原文摘要 · Abstract (English)

Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LLM-driven driving systems attach the language model as a post-processor and force it to reason over redundant or conflicting perception outputs, which can amplify hallucinated entities and unsafe conclusions. This paper proposes InfoCoordiBridge, a BEV-centric neuro-symbolic architecture that inserts an explicit coordination bridge between perception and language reasoning. InfoCoordiBridge comprises (i) a unified multi-agent perception layer that outputs typed structured facts together with modality-focused synopses, (ii) an ICA module that aligns and fuses multi-source outputs into a single SceneSummary, and (iii) an SSRE module that performs SceneSummary-grounded reasoning with verification. Experiments on nuScenes and Waymo show that ICA preserves competitive 3D detection accuracy while substantially improving fusion consistency, reducing redundancy to below 1% and achieving about 98% attribute agreement. On NuScenes-QA and a template-aligned Waymo-QA benchmark, SSRE improves factual grounding and reduces hallucinated entity mentions compared with representative VLM and agentic baselines. Overall, by coordinating multi-sensor outputs into a single conflict-aware SceneSummary before prompting, InfoCoordiBridge prevents redundant and cross-modally inconsistent perception evidence from propagating into high-level reasoning.

自动驾驶神经符号多模态融合幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。