arXiv:2601.04824cs.CV2026-01中稿 · Real World Surveil…

构建真实交通监控动作检索基准,评估模型对车辆行为的细粒度识别能力。

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

  • 基于监控视频设计车辆行为检索任务,区分相似动作并理解时间方向。
  • 现有视觉与多模态模型在细粒度动作区分上表现不佳,准确率低于人类。
  • 利用多模态大模型生成可解释描述,无需训练即实现高性能检索。

自动识别事件与重复行为分析对视频监控至关重要。然而,现有基于内容的视频检索基准多聚焦场景级相似性,未能评估监控中所需的动作区分能力。为此,我们提出SOVABench(Surveillance Opposite Vehicle Actions Benchmark),一个基于真实监控视频、聚焦车辆相关动作的检索基准。SOVABench定义了两种评估协议(跨对与同对),用于评估跨动作区分与时间方向理解能力。尽管动作区分对人类直观易懂,但实验表明,当前最先进的视觉与多模态模型仍难以处理。我们提出一种无需训练的框架,利用多模态大语言模型(MLLM)的视觉推理与指令遵循能力,从其生成的图文描述中提取可解释嵌入,该框架在SOVABench及多个空间与计数基准上表现优异,优于对比学习视觉-语言模型。代码、标注与构建指南已公开。

原文摘要 · Abstract (English)

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate the action discrimination required in surveillance. To address this gap, we introduce SOVABench (Surveillance Opposite Vehicle Actions Benchmark), a real-world retrieval benchmark built from surveillance footage and centered on vehicle-related actions. SOVABench defines two evaluation protocols (inter-pair and intra-pair) to assess cross-action discrimination and temporal direction understanding. Although action distinctions are generally intuitive for human observers, our experiments show that they remain challenging for state-of-the-art vision and multimodal models. Leveraging the visual reasoning and instruction-following capabilities of Multimodal Large Language Models (MLLMs), we present a training-free framework for producing interpretable embeddings from MLLM-generated descriptions for both images and videos. The framework achieves strong performance on SOVABench as well as on several spatial and counting benchmarks where contrastive Vision-Language Models often fail. The code, annotations, and instructions to construct the benchmark are publicly available.

视频检索多模态监控分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。