构建视频文本细粒度对齐新基准,提升模型理解连续事件能力
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
- 设计双基准测试,聚焦多事件视频与文本的时序对齐
- 引入细微时间扰动负样本,全面检验模型组合敏感性
- 提出分层配对偏好损失,适配少标注数据场景
我们提出 VideoComp,一个用于提升视觉语言模型在细粒度时序对齐方面能力的基准与学习框架。不同于以往关注静态图像或单一事件视频的基准,本工作聚焦连续多事件视频中的对齐问题。基于带有时间定位事件描述的数据集(如 ActivityNet-Captions、YouCook2),我们构建了两个组合性基准:ActivityNet-Comp 与 YouCook2-Comp。通过构造具有细微时间扰动的挑战性负样本(如事件重排序、动作词替换、部分描述、组合扰动),全面测试模型在长时序连贯视频-文本序列上的组合敏感性。为提升模型性能,我们提出一种分层配对偏好损失,强化与时间准确配对的对齐,并逐步惩罚逐渐被破坏的配对,促进细粒度组合学习。为缓解密集标注视频数据稀缺问题,我们引入一种预训练策略:将短视频-文本对拼接以模拟多事件序列。我们在 VideoComp 基准上评估了视频-文本基础模型与大型多模态模型(LMMs),揭示了现有模型在组合性理解方面的优劣。整体而言,本工作提供了一个全面的框架,用于评估与增强模型实现精细、时序一致的视频-文本对齐能力。
原文摘要 · Abstract (English)
We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event videos, our benchmark targets alignment in continuous multi-event videos. Leveraging video-text datasets with temporally localized event captions (e.g. ActivityNet-Captions, YouCook2), we construct two compositional benchmarks, ActivityNet-Comp and YouCook2-Comp. We create challenging negative samples with subtle temporal disruptions such as reordering, action word replacement, partial captioning, and combined disruptions. These benchmarks comprehensively test models' compositional sensitivity across extended, cohesive video-text sequences. To improve model performance, we propose a hierarchical pairwise preference loss that strengthens alignment with temporally accurate pairs and gradually penalizes increasingly disrupted ones, encouraging fine-grained compositional learning. To mitigate the limited availability of densely annotated video data, we introduce a pretraining strategy that concatenates short video-caption pairs to simulate multi-event sequences. We evaluate video-text foundational models and large multimodal models (LMMs) on our benchmark, identifying both strengths and areas for improvement in compositionality. Overall, our work provides a comprehensive framework for evaluating and enhancing model capabilities in achieving fine-grained, temporally coherent video-text alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。