arXiv:2512.01769cs.CVcs.DB2025-12

提出通用视频分析框架,自动识别跨领域复杂场景。

VideoScoop: A Non-Traditional Domain-Independent Framework For Video Analysis

  • 用关系模型和图模型统一表示视频内容
  • 支持连续查询与多场景检测,准确率超90%
  • 适合医疗监控、安防、辅助生活等多领域应用

自动理解视频内容对公共监测(CM)、一般监控(SL)、辅助生活(AL)等应用至关重要。尽管图像与视频分析(IVA)在内容提取(如目标识别与跟踪)方面取得进展,但识别有意义的活动或情境(如两物体靠近)仍具挑战,仅靠内容提取无法实现。当前视频情境分析(VSA)依赖人工或针对特定视频类型定制算法,缺乏通用性且需为每种新情境重新开发。本文提出一种通用型VSA框架,先通过先进视频内容提取技术一次性获取内容,再以扩展关系模型(R++)和图模型进行表示。使用R++可将内容作为数据流处理,支持通过提出的视频分析连续查询语言进行实时查询;图模型则能检测关系模型难以捕捉的情境。通过识别跨领域的基础情境变体并建模为参数化模板,实现领域无关性。在来自AL、CM、SL三个领域的不同长度视频数据集上,对多种情境进行了大量实验,验证了该方法在准确性、效率和鲁棒性方面的优越性。

原文摘要 · Abstract (English)

Automatically understanding video contents is important for several applications in Civic Monitoring (CM), general Surveillance (SL), Assisted Living (AL), etc. Decades of Image and Video Analysis (IVA) research have advanced tasks such as content extraction (e.g., object recognition and tracking). Identifying meaningful activities or situations (e.g., two objects coming closer) remains difficult and cannot be achieved by content extraction alone. Currently, Video Situation Analysis (VSA) is done manually with a human in the loop, which is error-prone and labor-intensive, or through custom algorithms designed for specific video types or situations. These algorithms are not general-purpose and require a new algorithm/software for each new situation or video from a new domain. This report proposes a general-purpose VSA framework that overcomes the above limitations. Video contents are extracted once using state-of-the-art Video Content Extraction technologies. They are represented using two alternative models -- the extended relational model (R++) and graph models. When represented using R++, the extracted contents can be used as data streams, enabling Continuous Query Processing via the proposed Continuous Query Language for Video Analysis. The graph models complement this by enabling the detection of situations that are difficult or impossible to detect using the relational model alone. Existing graph algorithms and newly developed algorithms support a wide variety of situation detection. To support domain independence, primitive situation variants across domains are identified and expressed as parameterized templates. Extensive experiments were conducted across several interesting situations from three domains -- AL, CM, and SL-- to evaluate the accuracy, efficiency, and robustness of the proposed approach using a dataset of videos of varying lengths from these domains.

视频分析情境检测通用框架图模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。