用多模态大模型提升交通感知与决策,让自动驾驶更安全可靠。
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
- 融合视觉、空间等多源数据,实现全局场景理解
- 显著增强对复杂路况的感知与抗干扰能力
- 适合智能驾驶、交通管理等系统研发人员参考
交通安全仍是全球重大挑战,传统高级驾驶辅助系统(ADAS)在动态真实场景中常因传感器数据割裂处理及易受对抗条件影响而失效。本文综述多模态大语言模型(MLLMs)在解决上述问题中的变革潜力,通过整合视觉、空间与环境等跨模态信息,实现对交通场景的全局理解。分析表明,基于MLLM的方法在感知、决策和对抗鲁棒性方面均有显著提升。关键数据集如KITTI、DRAMA、ML4RoadSafety推动了相关研究进展。未来方向包括实时边缘部署、因果推理驱动决策以及人机协同。本文强调,MLLMs可成为下一代交通安全部件的核心,提供可扩展、上下文感知的主动风险缓解方案,全面提升道路安全水平。
原文摘要 · Abstract (English)
Traffic safety remains a critical global challenge, with traditional Advanced Driver-Assistance Systems (ADAS) often struggling in dynamic real-world scenarios due to fragmented sensor processing and susceptibility to adversarial conditions. This paper reviews the transformative potential of Multimodal Large Language Models (MLLMs) in addressing these limitations by integrating cross-modal data such as visual, spatial, and environmental inputs to enable holistic scene understanding. Through a comprehensive analysis of MLLM-based approaches, we highlight their capabilities in enhancing perception, decision-making, and adversarial robustness, while also examining the role of key datasets (e.g., KITTI, DRAMA, ML4RoadSafety) in advancing research. Furthermore, we outline future directions, including real-time edge deployment, causality-driven reasoning, and human-AI collaboration. By positioning MLLMs as a cornerstone for next-generation traffic safety systems, this review underscores their potential to revolutionize the field, offering scalable, context-aware solutions that proactively mitigate risks and improve overall road safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。