arXiv:2507.04815cs.CVcs.AI2025-07被引 5

用时空事件图构建视觉与语言的可解释连接,自监督生成长篇视频描述。

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach

  • 基于时空事件图构建视觉与语言的共享表示,实现可解释的跨模态对齐。
  • 在多数据集上生成连贯、丰富且相关的长文本描述,超越简单枚举式摘要。
  • 可作为自监督教师模型,高效训练端到端神经学生模型,适合多任务理解研究者。

视频内容的自然语言描述任务常被称为视频字幕生成。与通常简短且广泛存在的传统视频字幕不同,以复杂语言形式生成的长篇段落描述极为稀缺。这一局限源于人工标注成本高,以及从底层故事角度解释语言形成过程的挑战性——即空间与时间中相互关联事件的复杂系统。通过分析近期方法与数据集,我们发现当前缺乏针对复杂语言描述视频的公开资源,远超简单字幕枚举的层次。尽管现有先进方法在直接端到端学习视频与文本间关系时能生成较短字幕,但解释视觉与语言间的深层关联仍难以实现。本文提出一种基于时空事件图的视觉-语言共享表示,可解释地整合多个视觉任务,生成最终自然语言描述。同时,我们证明该自动化可解释描述生成流程可作为自监督教师,有效训练直接端到端神经学生路径,构建神经-分析混合系统。在多个多样化数据集上验证,该方法在标准评估指标、人工标注及多个先进视觉语言模型共识下,均能生成连贯、丰富且相关性强的文本描述。

原文摘要 · Abstract (English)

The task of describing video content in natural language is commonly referred to as video captioning. Unlike conventional video captions, which are typically brief and widely available, long-form paragraph descriptions in natural language are scarce. This limitation of current datasets is due to the expensive human manual annotation required and to the highly challenging task of explaining the language formation process from the perspective of the underlying story, as a complex system of interconnected events in space and time. Through a thorough analysis of recently published methods and available datasets, we identify a general lack of published resources dedicated to the problem of describing videos in complex language, beyond the level of descriptions in the form of enumerations of simple captions. Furthermore, while state-of-the-art methods produce impressive results on the task of generating shorter captions from videos by direct end-to-end learning between the videos and text, the problem of explaining the relationship between vision and language is still beyond our reach. In this work, we propose a shared representation between vision and language, based on graphs of events in space and time, which can be obtained in an explainable and analytical way, to integrate and connect multiple vision tasks to produce the final natural language description. Moreover, we also demonstrate how our automated and explainable video description generation process can function as a fully automatic teacher to effectively train direct, end-to-end neural student pathways, within a self-supervised neuro-analytical system. We validate that our explainable neuro-analytical approach generates coherent, rich and relevant textual descriptions on videos collected from multiple varied datasets, using both standard evaluation metrics, human annotations and consensus from ensembles of state-of-the-art VLMs.

视频描述可解释性自监督事件图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。