arXiv:2412.07704cs.CV2024-12被引 2

解决视频语言模型多粒度对齐难题,支持长视频理解。

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning

  • 通过数据扩展与迭代逼近模块,统一建模多粒度视频文本。
  • 在7个数据集上达到领先或相当水平,长视频任务表现突出。
  • 适用于任意粒度数量,可扩展性强,适合多粒度视频任务研究者。

在各类视频-语言学习任务中,实现多粒度数据的跨模态对齐仍具挑战。本文从数据和建模两方面提出解决方案:针对缺乏多粒度视频-文本预训练数据的问题,提出粒度扩展(GEX)方法,通过集成与压缩操作扩展单粒度数据的粒度;为更好建模多粒度数据,引入迭代逼近模块(IAM),将多粒度视频与文本嵌入统一的低维语义空间,同时保留跨模态对齐所需的关键信息。GEXIA具有高度可扩展性,不限制对齐的粒度数量。我们在三个视频任务类别下的七个基准数据集上进行评估,展现出先进或相当的性能。尤为显著的是,尽管预训练数据仅包含短视频片段,该模型在长视频理解任务中仍表现出色。

原文摘要 · Abstract (English)

In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from two crucial perspectives: data and modeling. Given the absence of a multi-grained video-text pretraining dataset, we introduce a Granularity EXpansion (GEX) method with Integration and Compression operations to expand the granularity of a single-grained dataset. To better model multi-grained data, we introduce an Iterative Approximation Module (IAM), which embeds multi-grained videos and texts into a unified, low-dimensional semantic space while preserving essential information for cross-modal alignment. Furthermore, GEXIA is highly scalable with no restrictions on the number of video-text granularities for alignment. We evaluate our work on three categories of video tasks across seven benchmark datasets, showcasing state-of-the-art or comparable performance. Remarkably, our model excels in tasks involving long-form video understanding, even though the pretraining dataset only contains short video clips.

视频语言多粒度跨模态对齐长视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。