arXiv:2606.12047cs.CVcs.AI2026-06中稿 · CVPR

通过分阶段推理,让AI从监控视频中零样本理解事故的时间、类型和位置。

Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding

论文配图:Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
图 1 · 摘自论文原文
  • 分三步:先定位时间窗,再多视角提示推理,最后空间定位事故点。
  • 在零样本基准上,调和平均得分显著优于中心点基线。
  • 适合需要可靠视觉语言推理的自动驾驶与安防场景。

本文针对从监控视频中零样本理解事故的问题,通过自然语言识别事故发生的时间、类型及在画面中的位置。提出三阶段流程:第一阶段利用视觉-语言相似性提取事故附近短时窗;第二阶段采用五种互补视角(基础、运动、几何、对比、仲裁)进行元数据驱动的多提示推理,并通过熵门控成对裁决解决分歧;第三阶段使用开放词汇检测器查询预测的事故类型与场景布局,结合关键帧检测结果,以得分加权质心聚合定位。该流程在零样本ACCIDENT@CVPR基准上,相较于中心点基线显著提升调和平均得分。实验表明,将零样本视频理解分解为时间定位、语义分类与空间定位,可使视觉语言模型实现更可靠的推理。

原文摘要 · Abstract (English)

In this paper, we address the problem of zero-shot understanding of accidents from surveillance videos by identifying when an impact event occurs, what type of impact it is, and where in the frame it occurs using natural language. We propose a three-stage pipeline that decomposes the accident understanding into when, what, and where. The first stage extracts a short temporal window around the impact using vision-language similarity. In the second stage, we perform metadata-driven multi-prompt reasoning with five complementary views (baseline, motion, geometry, contrast, and tiebreaker) and resolve disagreement via an entropy-gated pairwise adjudicator. Finally, we localize the impact of an open-vocabulary detector queried on the predicted accident type and scene layout, and aggregate detections across keyframes using a score-weighted centroid. Our pipeline achieves a substantial improvement in the harmonic-mean score over a centre-of-frame baseline on the zero-shot ACCIDENT @ CVPR benchmark. We show that decomposing zero-shot video understanding into temporal localization, semantic classification, and spatial grounding enable more reliable reasoning with vision-language models than direct prompting alone.

零样本事故理解视觉语言多提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。