arXiv:2508.06763cs.CVcs.AI2025-08被引 9

让多模态大模型看懂交通事故的像素级细节和时间顺序

SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding

  • 引入像素级理解与时间定位能力,支持区域问答和语言驱动分割
  • 在多个任务上表现优异,尤其在细粒度场景分析中显著超越现有模型
  • 适合智能交通、自动驾驶安全分析等需要精准事故理解的场景

多模态大语言模型(MLLM)在视觉-语言任务中取得显著进展,并展现出在交通事故理解方面的巨大潜力。然而,现有方法主要依赖粗粒度图像或视频层面的理解,难以处理细粒度视觉细节或局部场景组件,限制了其在复杂事故场景中的应用。为此,我们提出SafePLUG框架,赋予MLLM像素级理解与时间定位能力,实现对交通事故的全面分析。该框架支持任意形状视觉提示下的区域感知问答、基于语言指令的像素级分割,以及事故中时序锚定事件的识别。为推动该领域发展,我们构建了一个新数据集,包含多样化事故场景的多模态问答对,附带详细的像素级标注与时间事件边界。实验表明,SafePLUG在区域问答、像素分割、时间事件定位及事故理解等多项任务中均表现强劲。这些能力为复杂交通场景的细粒度理解奠定基础,有望提升驾驶安全与智慧交通系统的态势感知能力。代码、数据集与模型权重将公开发布于:https://zihaosheng.github.io/SafePLUG

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved remarkable progress across a range of vision-language tasks and demonstrate strong potential for traffic accident understanding. However, existing MLLMs in this domain primarily focus on coarse-grained image-level or video-level comprehension and often struggle to handle fine-grained visual details or localized scene components, limiting their applicability in complex accident scenarios. To address these limitations, we propose SafePLUG, a novel framework that empowers MLLMs with both Pixel-Level Understanding and temporal Grounding for comprehensive traffic accident analysis. SafePLUG supports both arbitrary-shaped visual prompts for region-aware question answering and pixel-level segmentation based on language instructions, while also enabling the recognition of temporally anchored events in traffic accident scenarios. To advance the development of MLLMs for traffic accident understanding, we curate a new dataset containing multimodal question-answer pairs centered on diverse accident scenarios, with detailed pixel-level annotations and temporal event boundaries. Experimental results show that SafePLUG achieves strong performance on multiple tasks, including region-based question answering, pixel-level segmentation, temporal event localization, and accident event understanding. These capabilities lay a foundation for fine-grained understanding of complex traffic scenes, with the potential to improve driving safety and enhance situational awareness in smart transportation systems. The code, dataset, and model checkpoints will be made publicly available at: https://zihaosheng.github.io/SafePLUG

多模态模型交通事故像素理解时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。