通过轨迹引导的像素时间对齐,提升视频语言模型的细粒度理解能力。
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
- 基于物体运动轨迹实现视觉与语言在时空维度的精细对齐。
- 在143,000个视频-文本对上训练,显著超越现有方法。
- 适合需要精准视频语义理解的研究者与开发者。
受大语言模型(LLMs)浪潮推动,大型视觉语言模型(LVLMs)已成为连接图像与文本的关键进展。然而,视频由于其时空结构复杂,使得LVLM难以有效处理。现有大型视频语言模型(LVidLMs)通常将静态视觉数据(如图像)特征对齐到语言特征的潜在空间,通过通用多模态任务充分挖掘LLM的能力。本文提出一种细粒度对齐方法,利用物体轨迹在空间与时间维度上同步对齐不同模态。为此,我们构建了名为PiTe-143k的多模态预训练数据集,通过自动标注流程为视频和字幕中出现的所有对象提供像素级运动轨迹。基于此,我们提出新的LVidLM——PiTe,展现出优异的可应用性,在多种视频相关多模态任务中大幅超越现有最先进方法。
原文摘要 · Abstract (English)
Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and spatial-temporal data structure. Recent Large Video-Language Models (LVidLMs) align feature of static visual data like image into latent space of language feature, by general multi-modal tasks to leverage abilities of LLMs sufficiently. In this paper, we explore fine-grained alignment approach via object trajectory for different modalities across both spatial and temporal dimensions simultaneously. Thus, we propose a novel LVidLM by trajectory-guided Pixel-Temporal Alignment, dubbed PiTe, that exhibits promising applicable model property. To achieve fine-grained video-language alignment, we curate a multi-modal pre-training dataset PiTe-143k, the dataset provision of moving trajectories in pixel level for all individual objects, that appear and mention in the video and caption both, by our automatic annotation pipeline. Meanwhile, PiTe demonstrates astounding capabilities on myriad video-related multi-modal tasks through beat the state-of-the-art methods by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。