用时空事件图推理实现零样本可解释视频描述
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
- 构建时空事件图表示视频内容,统一视觉与语言理解
- 在多个数据集上生成连贯丰富的自然语言描述,超越传统方法
- 适合需要透明化决策过程的视频理解应用
当前机器学习中,Transformer已成为计算机视觉与自然语言处理等领域的主流方法。尽管基于Transformer的方法在图像、视频分类、分割、动作识别等方面表现优异,但视觉与语言之间的关系理解仍面临挑战。本文提出一种基于时空事件图的可解释、程序化方法,为视觉与语言建立共同语义基础,连接现有学习型视觉与语言模型,解决视频自然语言描述这一长期难题。实验表明,该算法在多种数据集上均能生成连贯、丰富且相关的文本描述,使用标准指标(如Bleu、ROUGE)和现代大模型作为评判者(LLM-as-a-Jury)进行验证,效果显著。
原文摘要 · Abstract (English)
In the current era of Machine Learning, Transformers have become the de facto approach across a variety of domains, such as computer vision and natural language processing. Transformer-based solutions are the backbone of current state-of-the-art methods for language generation, image and video classification, segmentation, action and object recognition, among many others. Interestingly enough, while these state-of-the-art methods produce impressive results in their respective domains, the problem of understanding the relationship between vision and language is still beyond our reach. In this work, we propose a common ground between vision and language based on events in space and time in an explainable and programmatic way, to connect learning-based vision and language state of the art models and provide a solution to the long standing problem of describing videos in natural language. We validate that our algorithmic approach is able to generate coherent, rich and relevant textual descriptions on videos collected from a variety of datasets, using both standard metrics (e.g. Bleu, ROUGE) and the modern LLM-as-a-Jury approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。