通过反转视频构造更难负样本,提升视频文本检索的时序理解挑战。
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
- 用反转真实动作视频生成难负例,强化时序建模能力评估。
- 构建含21000个视频、每视频10条描述的RTime数据集,总时长122小时。
- 提出三类新基准任务,推动视频-文本模型在时序理解上的进步。
跨模态检索是信息检索与多模态视觉语言理解的重要任务。视频-文本检索因需理解时间动态而比图像-文本检索更具挑战性。然而,我们发现现有视频-文本基准在全面评估模型能力方面存在不足,尤其在时序理解上,导致大规模图像-文本预训练模型已能实现与视频-文本预训练模型相当的零样本性能。本文提出RTime,一个强调时序特性的新型视频-文本检索数据集。我们首先选取具有显著时序特征的动作或事件视频,然后反转这些视频以生成更难的负样本。随后,招募标注者判断候选视频的时序显著性与可逆性,并撰写相应标题。进一步利用GPT-4基于人工标题扩展更多标题。当前RTime数据集包含21,000个视频,每个视频有10条标题,总计约122小时。基于RTime,我们设计三个检索基准任务:RTime-Origin、RTime-Hard和RTime-Binary。同时,在模型训练中引入更难负样本,对多种视频-文本模型进行基准测试。大量实验分析表明,RTime确实为视频-文本检索带来了新的更高挑战。我们已公开RTime数据集(https://github.com/qyr0403/Reversed-in-Time),以推动视频-文本检索与多模态理解研究发展。
原文摘要 · Abstract (English)
Cross-modal (e.g. image-text, video-text) retrieval is an important task in information retrieval and multimodal vision-language understanding field. Temporal understanding makes video-text retrieval more challenging than image-text retrieval. However, we find that the widely used video-text benchmarks have shortcomings in comprehensively assessing abilities of models, especially in temporal understanding, causing large-scale image-text pre-trained models can already achieve comparable zero-shot performance with video-text pre-trained models. In this paper, we introduce RTime, a novel temporal-emphasized video-text retrieval dataset. We first obtain videos of actions or events with significant temporality, and then reverse these videos to create harder negative samples. We then recruit annotators to judge the significance and reversibility of candidate videos, and write captions for qualified videos. We further adopt GPT-4 to extend more captions based on human-written captions. Our RTime dataset currently consists of 21k videos with 10 captions per video, totalling about 122 hours. Based on RTime, we propose three retrieval benchmark tasks: RTime-Origin, RTime-Hard, and RTime-Binary. We further enhance the use of harder-negatives in model training, and benchmark a variety of video-text models on RTime. Extensive experiment analysis proves that RTime indeed poses new and higher challenges to video-text retrieval. We release our RTime dataset https://github.com/qyr0403/Reversed-in-Time to further advance video-text retrieval and multimodal understanding research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。