MMDrive融合多模态数据,实现3D交通场景的深度理解。
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- 引入占据图、点云与文本,实现跨模态动态融合。
- 在DriveLM上达BLEU-4 54.56、METEOR 41.78,NuScenes-QA准确率62.7%。
- 适合关注自动驾驶场景理解与可解释性的研究者。
视觉语言模型通过多源信息融合实现复杂交通场景的理解与推理,已成为自动驾驶核心技术。然而,现有模型受限于二维平面图像理解范式,难以感知三维空间信息并进行深层语义融合,导致在复杂驾驶环境中表现不佳。本研究提出MMDrive,一种将传统图像理解扩展为通用3D场景理解框架的多模态视觉语言模型。MMDrive融合占据图、激光雷达点云和文本场景描述三种互补模态,提出两个新组件:文本导向的多模态调制器根据问题语义动态加权各模态贡献,实现上下文感知特征融合;跨模态抽象器使用可学习抽象令牌生成紧凑的跨模态摘要,突出关键区域与核心语义。在DriveLM和NuScenes-QA基准上的全面评估表明,MMDrive显著优于现有自动驾驶视觉语言模型,在DriveLM上取得BLEU-4 54.56和METEOR 41.78,在NuScenes-QA上达到62.7%准确率。MMDrive有效突破传统图像单模理解瓶颈,实现在复杂驾驶环境中的鲁棒多模态推理,为可解释的自动驾驶场景理解提供新基础。
原文摘要 · Abstract (English)
Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are constrained by the image understanding paradigm in 2D plane, which restricts their capability to perceive 3D spatial information and perform deep semantic fusion, resulting in suboptimal performance in complex autonomous driving environments. This study proposes MMDrive, an multimodal vision-language model framework that extends traditional image understanding to a generalized 3D scene understanding framework. MMDrive incorporates three complementary modalities, including occupancy maps, LiDAR point clouds, and textual scene descriptions. To this end, it introduces two novel components for adaptive cross-modal fusion and key information extraction. Specifically, the Text-oriented Multimodal Modulator dynamically weights the contributions of each modality based on the semantic cues in the question, guiding context-aware feature integration. The Cross-Modal Abstractor employs learnable abstract tokens to generate compact, cross-modal summaries that highlight key regions and essential semantics. Comprehensive evaluations on the DriveLM and NuScenes-QA benchmarks demonstrate that MMDrive achieves significant performance gains over existing vision-language models for autonomous driving, with a BLEU-4 score of 54.56 and METEOR of 41.78 on DriveLM, and an accuracy score of 62.7% on NuScenes-QA. MMDrive effectively breaks the traditional image-only understanding barrier, enabling robust multimodal reasoning in complex driving environments and providing a new foundation for interpretable autonomous driving scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。