用知识图谱替代自然语言描述视频,提升可计算性和准确性
Detection-Fusion for Knowledge Graph Extraction from Videos
- 先预测视频中人物对,再识别其关系,构建结构化知识图谱
- 引入外部背景知识,增强知识图谱的语义丰富性与一致性
- 适合需要精准理解视频内容的智能系统开发者
视频理解中的一个挑战性任务是从视频输入中提取语义内容。现有系统多依赖语言模型生成自然语言描述,但存在过度依赖语言模型、输出基于文本统计规律而非视觉内容的问题。此外,自然语言标注难以计算机处理,评估困难,且不易翻译。本文提出一种新方法,通过知识图谱标注视频,克服上述问题。具体地,设计了一种基于深度学习的模型,先预测视频中的人物对,再识别其关系。同时,提出模型扩展以整合背景知识,辅助知识图谱构建。
原文摘要 · Abstract (English)
One of the challenging tasks in the field of video understanding is extracting semantic content from video inputs. Most existing systems use language models to describe videos in natural language sentences, but this has several major shortcomings. Such systems can rely too heavily on the language model component and base their output on statistical regularities in natural language text rather than on the visual contents of the video. Additionally, natural language annotations cannot be readily processed by a computer, are difficult to evaluate with performance metrics and cannot be easily translated into a different natural language. In this paper, we propose a method to annotate videos with knowledge graphs, and so avoid these problems. Specifically, we propose a deep-learning-based model for this task that first predicts pairs of individuals and then the relations between them. Additionally, we propose an extension of our model for the inclusion of background knowledge in the construction of knowledge graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。