arXiv:2606.02443cs.CLcs.AI2026-06被引 1

构建首个流式视频安全预警基准,测试模型实时发现风险的能力。

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

论文配图:PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
图 1 · 摘自论文原文
  • 设计740个视频的流式基准,标注每段视频的风险起始与事故边界。
  • 13个大模型在严格指标下最高仅20.0%准确率,高召回必伴高误报。
  • 模型在日常场景表现较好,驾驶场景误报严重,依赖表面动作而非深层推理。

从危险初现到事故发生之间,通常存在可干预的时间窗口。具备视频理解能力的多模态大语言模型(MLLMs)可作为持续监控系统,在此窗口内发出预警。然而现有基准无法测试该能力:它们依赖静态输入,忽略时间精度,且未评估安全场景中的误报率。本文提出PaSBench-Video,一个包含740个视频的基准,涵盖驾驶、医疗、日常生活和工业生产四个领域,其中481个为风险视频,259个为无风险视频。风险视频均标注帧级风险起始点和事故边界。模型需因果地观察视频,并生成时间精准且内容正确的预警。测试13个MLLMs后发现,无一模型在最严格指标下超过20.0%得分;召回率与误报率呈显著正相关(Pearson相关系数0.64),即提升检测率必然导致大量安全片段误报。性能在不同领域差异明显:日常场景中模型可在低误报下实现中等召回,因风险具有异常性;而在驾驶场景中,模型几乎对所有视频误报,因常规与危险场景外观相似。结果表明,当前模型依赖场景级活动线索,而非对潜在危害的推理。

原文摘要 · Abstract (English)

Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable multimodal large language models (MLLMs) could serve as always-on safety monitors that issue warnings during this window. Yet current benchmarks do not test this ability: they rely on static inputs, ignore timing precision, and omit false-positive measurement on safe scenes. We present PaSBench-Video, a 740-video benchmark with 481 risk and 259 no-risk videos across four domains: driving, healthcare, daily life, and industrial production. Risk videos are annotated with frame-level risk onset and accident boundaries. A model must observe the video causally and produce a warning that is both temporally calibrated and content-correct. Testing 13 MLLMs, we find that no model exceeds 20.0% on our strictest metric, and recall is tightly coupled with false-positive rate, with Pearson correlation 0.64: higher detection comes only at the cost of triggering warnings on the majority of safe clips. Performance splits sharply by domain: models achieve moderate recall at low false-positive rates in daily life, where risks are inherently anomalous, yet fire indiscriminately in driving, where routine and hazardous scenes look alike. These results indicate that current models rely on scene-level activity cues rather than reasoning about emerging harm.

视频预警多模态安全监测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。