arXiv:2603.06576cs.CVcs.AI2026-03中稿 · ECCV

将大模型语义知识蒸馏到鸟瞰图表征,提升自动驾驶跨视角推理能力。

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

  • 用鸟瞰图统一多视角输入,实现空间一致的视觉表示。
  • 在复杂场景下使大模型推理准确率提升46.0%,闭环驾驶性能最高提升28.2%。
  • 适合需要强语义理解与几何一致性融合的自动驾驶研究者。

将大型语言模型(LLMs)融入自动驾驶系统,可增强其在复杂决策和长尾场景下的推理与语义理解能力。然而,现有方法通常独立处理多视角、多帧图像的文本标记,导致计算冗余且空间一致性差。这种视觉处理的分离限制了三维空间推理的准确性,并破坏了视图间的几何连贯性。另一方面,基于几何标注任务(如目标检测)学习的鸟瞰图(BEV)表征虽具备空间结构,但缺乏基础视觉编码器的语义丰富性。为此,我们提出BEVLM框架,将空间一致且语义蒸馏的BEV表征与LLMs相连接。大量实验表明,通过使用BEV特征作为统一输入,BEVLM显著提升了LLMs在跨视角驾驶场景中的推理效率,准确率提高46.0%。此外,通过将大模型的语义知识蒸馏至BEV表征,该方法在UniAD和VAD数据集上,于安全关键场景中实现了最高达28.2%的端到端驾驶性能提升。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios. However, existing methods typically feed LLMs with tokens from multi-view and multi-frame images independently, leading to redundant computation and limited spatial consistency. This separation in visual processing hinders accurate 3D spatial reasoning and fails to maintain geometric coherence across views. On the other hand, Bird's-Eye View (BEV) representations learned from geometrically annotated tasks (e.g., object detection) provide spatial structure but lack the semantic richness of foundation vision encoders. To bridge this gap, we propose BEVLM, a framework that connects a spatially consistent and semantically distilled BEV representation with LLMs. Through extensive experiments, we show that BEVLM enables LLMs to reason more effectively in cross-view driving scenes, improving accuracy by 46.0%, by leveraging BEV features as unified inputs. Furthermore, by distilling semantic knowledge from LLMs into BEV representations, BEVLM significantly improves closed-loop end-to-end driving performance in safety-critical scenarios across UniAD and VAD, with gains of up to 28.2%.

自动驾驶鸟瞰图大模型知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。