将视频转化为可查询的知识图谱,支持持续学习与领域知识更新。
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
- 融合预训练模型,将视频转为时序半结构化数据。
- 构建帧级可查询的知识图谱,支持动态添加新知识。
- 适合需要持续更新知识的多模态分析场景。
多模态内容分析往往计算复杂且工程成本高。尽管已有大量针对静态数据的预训练模型,但将其与视频等复杂数据融合仍具挑战。本文提出一个框架,可高效原型化多模态内容分析流程。我们设计了一种管道方案,整合一组预训练模型,将视频转换为时序半结构化数据格式,并进一步转化为帧级索引知识图谱。该表示支持查询与持续学习,可通过交互式方式动态融入新领域知识。
原文摘要 · Abstract (English)
Analysis of multi-modal content can be tricky, computationally expensive, and require a significant amount of engineering efforts. Lots of work with pre-trained models on static data is out there, yet fusing these opensource models and methods with complex data such as videos is relatively challenging. In this paper, we present a framework that enables efficiently prototyping pipelines for multi-modal content analysis. We craft a candidate recipe for a pipeline, marrying a set of pre-trained models, to convert videos into a temporal semi-structured data format. We translate this structure further to a frame-level indexed knowledge graph representation that is query-able and supports continual learning, enabling the dynamic incorporation of new domain-specific knowledge through an interactive medium.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。