用大模型提升交通场景图像分割,让自动驾驶更懂路况。
Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems

- 结合大模型提示技术增强图像分割能力
- 提升自动驾驶与交通监控的场景理解精度
- 适合关注智能交通AI落地的研究者与工程师
大型语言模型(LLMs)与计算机视觉的融合正深刻改变图像分割等感知任务。在智能交通系统(ITS)中,精准的场景理解对安全与效率至关重要,这一新范式提供了前所未有的能力。本综述系统梳理了大模型增强图像分割的新兴领域,聚焦其在ITS中的应用、挑战与未来方向。我们基于提示机制与核心架构提出分类体系,阐明这些创新如何提升自动驾驶、交通监测与基础设施维护中的道路场景理解能力。最后,我们识别出实时性能与安全关键可靠性等关键挑战,并提出以可解释、以人为本的AI为前提,推动该技术在下一代交通系统中的成功部署。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) with computer vision is profoundly transforming perception tasks like image segmentation. For intelligent transportation systems (ITS), where accurate scene understanding is critical for safety and efficiency, this new paradigm offers unprecedented capabilities. This survey systematically reviews the emerging field of LLM-augmented image segmentation, focusing on its applications, challenges, and future directions within ITS. We provide a taxonomy of current approaches based on their prompting mechanisms and core architectures, and we highlight how these innovations can enhance road scene understanding for autonomous driving, traffic monitoring, and infrastructure maintenance. Finally, we identify key challenges, including real-time performance and safety-critical reliability, and outline a perspective centered on explainable, human-centric AI as a prerequisite for the successful deployment of this technology in next-generation transportation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。