让自动驾驶大模型仅靠2D图像就能精准感知空间,无需3D模型
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

- 通过相机位姿注入和多视角对极注意力,增强2D图像间的空间对应关系
- 在nuScenes等5个基准上超越现有方法,风险检测准确率提升8.3%
- 适合需要轻量级空间推理的自动驾驶系统研发人员
多模态大语言模型(MLLM)在2D视觉任务中表现优异,但在自动驾驶等真实场景中的空间智能仍面临挑战。现有几何感知方法依赖推理时的辅助3D模型,导致流程复杂且易出错。本文提出OmniSpace,一种仅基于2D观测的即插即用几何感知范式。针对当前模型在跨视角对应与深度估计方面的瓶颈,OmniSpace引入相机位姿注入、多视角对极注意力模块及3D几何蒸馏目标,联合提升空间理解能力。大量实验表明,OmniSpace在规划任务(nuScenes、Bench2Drive)、风险检测(nuInstruct)、语言理解(Omnidrive)和泛化性(DriveBench)上均优于现有方法,显著提升性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-play paradigm for geometry-aware spatial reasoning from purely 2D observations. Motivated by our finding that current MLLMs are bottlenecked by weak cross-view correspondence and depth estimation, OmniSpace introduces a Camera Pose Injector, a Multi-view Epipolar Attention module, and a 3D Geometric Distillation objective that jointly address these two limitations by transferring geometric knowledge into the model. Extensive experiments show that OmniSpace surpasses existing methods on planning benchmarks (nuScenes, Bench2Drive), risk detection (nuInstruct), language (Omnidrive), and generalization (DriveBench).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。