提出深度安全视频理解新任务,可分析威胁成因并量化其严重性。
DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE
- 构建统一物理世界正则化MoE模型,融合多层级现实信息
- 在UCF-C与CUVA数据集上超越现有视频大模型性能
- 适合安全监控、事件溯源等需要深层推理的场景
现有安全视频理解研究多聚焦于检测和定位威胁(如枪击、抢劫),但缺乏对威胁成因的有效生成与评估能力。为此,本文提出新的深度安全视频理解(DeepSVU)任务,不仅识别威胁位置,还追溯并评估其成因。该任务面临两大挑战:如何有效建模从粗到细的物理世界信息(如人类行为、物体交互与背景上下文);如何自适应权衡这些因素。为此,本文提出统一物理世界正则化门控专家模型(UPRM),包含统一物理增强门控专家(UPE)模块与物理世界权衡正则化器(PTR),分别应对上述挑战。在自建的DeepSVU指令数据集UCF-C Instructions和CUVA Instructions上的大量实验表明,UPRM优于多个先进视频大模型及非VLM方法,验证了细粒度物理世界信息的重要性及其在模型中的有效捕捉能力。
原文摘要 · Abstract (English)
In the literature, prior research on Security-oriented Video Understanding (SVU) has predominantly focused on detecting and localize the threats (e.g., shootings, robberies) in videos, while largely lacking the effective capability to generate and evaluate the threat causes. Motivated by these gaps, this paper introduces a new chat paradigm SVU task, i.e., In-depth Security-oriented Video Understanding (DeepSVU), which aims to not only identify and locate the threats but also attribute and evaluate the causes threatening segments. Furthermore, this paper reveals two key challenges in the proposed task: 1) how to effectively model the coarse-to-fine physical-world information (e.g., human behavior, object interactions and background context) to boost the DeepSVU task; and 2) how to adaptively trade off these factors. To tackle these challenges, this paper proposes a new Unified Physical-world Regularized MoE (UPRM) approach. Specifically, UPRM incorporates two key components: the Unified Physical-world Enhanced MoE (UPE) Block and the Physical-world Trade-off Regularizer (PTR), to address the above two challenges, respectively. Extensive experiments conduct on our DeepSVU instructions datasets (i.e., UCF-C instructions and CUVA instructions) demonstrate that UPRM outperforms several advanced Video-LLMs as well as non-VLM approaches. Such information.These justify the importance of the coarse-to-fine physical-world information in the DeepSVU task and demonstrate the effectiveness of our UPRM in capturing such information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。