提升视频CLIP模型对长描述的理解能力,让机器更懂视频细节。
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
- 构建大规模视频-长描述配对数据集,支持长文本训练。
- 提出TPCM方法,增强特征空间分布学习,提升长描述理解。
- 新增两个评估任务,适合研究视频细节与幻觉识别的学者。
对比语言-图像预训练(CLIP)在众多应用中广受关注,但其预训练阶段侧重简短摘要文本,导致难以理解长描述。这一问题在视频场景尤为突出,因视频常包含丰富细节。本文提出VideoCLIP-XL(eXtra Length)模型,旨在释放视频CLIP模型对长描述的理解潜力。首先,我们建立自动化数据收集系统,构建大规模视频-长描述配对的VILD预训练数据集。其次,提出文本相似性引导的主要成分匹配(TPCM)方法,以更好学习特征空间分布并扩展长描述能力。同时引入两个新任务:细节感知描述排序(DDR)和幻觉感知描述排序(HDR),进一步提升理解性能。最后,构建长视频描述排序(LVDR)基准,实现更全面的评估。在包含短/长描述的多个主流文本-视频检索基准及我们自建的LVDR基准上,实验结果充分证明了该方法的有效性。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions. This issue is particularly acute regarding videos given that videos often contain abundant detailed contents. In this paper, we propose the VideoCLIP-XL (eXtra Length) model, which aims to unleash the long-description understanding capability of video CLIP models. Firstly, we establish an automatic data collection system and gather a large-scale VILD pre-training dataset with VIdeo and Long-Description pairs. Then, we propose Text-similarity-guided Primary Component Matching (TPCM) to better learn the distribution of feature space while expanding the long description capability. We also introduce two new tasks namely Detail-aware Description Ranking (DDR) and Hallucination-aware Description Ranking (HDR) for further understanding improvement. Finally, we construct a Long Video Description Ranking (LVDR) benchmark for evaluating the long-description capability more comprehensively. Extensive experimental results on widely-used text-video retrieval benchmarks with both short and long descriptions and our LVDR benchmark can fully demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。