综述自动驾驶中多模态感知与大模型融合的最新进展
All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles

- 系统梳理摄像头、激光雷达等传感器及其融合策略
- 提出面向车联环境的数据集分类框架,涵盖V2X等新型数据结构
- 聚焦视觉语言模型与大模型在目标检测中的应用前景
自动驾驶汽车正通过智能感知、决策与控制系统重塑交通未来。其成功依赖于复杂多模态环境下可靠的目标检测能力。尽管计算机视觉与人工智能取得显著进展,知识仍分散于多模态感知、上下文推理与协同智能之间。本文综述填补该空白,聚焦生成式AI、视觉语言模型(VLMs)与大型语言模型(LLMs)等新兴范式,而非回顾过时技术。首先系统分析车载传感器(相机、超声波、激光雷达、雷达)及其融合策略,揭示其在动态驾驶环境中的优劣,并探讨与LLM/VLM驱动感知框架的集成潜力。接着提出结构化分类的自动驾驶数据集体系,涵盖自车、基础设施及协作式数据集(如V2V、V2I、V2X、I2I),并进行跨数据结构对比分析。最后深入解析前沿检测方法,包括2D/3D流水线、混合传感器融合,特别关注基于视觉变压器(ViTs)、大/小语言模型(SLMs)与VLMs的新型Transformer架构。通过整合多维视角,本综述明确当前能力边界、开放挑战与未来方向。
原文摘要 · Abstract (English)
Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control systems. However, their success is tied to one core capability, reliable object detection in complex and multimodal environments. While recent breakthroughs in Computer Vision (CV) and Artificial Intelligence (AI) have driven remarkable progress, the field still faces a critical challenge as knowledge remains fragmented across multimodal perception, contextual reasoning, and cooperative intelligence. This survey bridges that gap by delivering a forward-looking analysis of object detection in AVs, emphasizing emerging paradigms such as Vision-Language Models (VLMs), Large Language Models (LLMs), and Generative AI rather than re-examining outdated techniques. We begin by systematically reviewing the fundamental spectrum of AV sensors (camera, ultrasonic, LiDAR, and Radar) and their fusion strategies, highlighting not only their capabilities and limitations in dynamic driving environments but also their potential to integrate with recent advances in LLM/VLM-driven perception frameworks. Next, we introduce a structured categorization of AV datasets that moves beyond simple collections, positioning ego-vehicle, infrastructure-based, and cooperative datasets (e.g., V2V, V2I, V2X, I2I), followed by a cross-analysis of data structures and characteristics. Ultimately, we analyze cutting-edge detection methodologies, ranging from 2D and 3D pipelines to hybrid sensor fusion, with particular attention to emerging transformer-driven approaches powered by Vision Transformers (ViTs), Large and Small Language Models (SLMs), and VLMs. By synthesizing these perspectives, our survey delivers a clear roadmap of current capabilities, open challenges, and future opportunities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。