构建细粒度视频定位新基准,提升精准视频片段检索能力。
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

- 用大模型自动生成带精细静态与动态描述的视频字幕
- 通过噪声评估器过滤幻觉内容,确保标注高质量
- 新基准涵盖多个数据集,适合研究细粒度视频理解
现有视频语料段落检索(VCMR)局限于粗粒度理解,难以应对细粒度查询下的精确定位。本文提出更挑战性的细粒度VCMR基准,要求模型从语料库中定位最佳匹配片段,需区分部分匹配候选。为提高数据构建效率并保证标注质量,我们提出VERIFIED——一个基于大语言模型(LLM)与多模态模型(LMM)的自动视频文本标注流水线,包含静态与动态增强字幕生成模块。为消除LLM幻觉导致的错误标注,我们设计细粒度感知噪声评估器,通过扰动硬负样本与对比损失微调视频基础模型。基于VERIFIED,我们构建了包含Charades-FIG、DiDeMo-FIG、ActivityNet-FIG的细粒度VCMR基准,标注质量高。在该数据集上评估多个先进VCMR模型,结果表明细粒度视频理解仍有巨大提升空间。代码与数据已开源。
原文摘要 · Abstract (English)
Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding, which hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from the corpus with other partially matched candidates. To improve the dataset construction efficiency and guarantee high-quality data annotations, we propose VERIFIED, an automatic \underline{V}id\underline{E}o-text annotation pipeline to generate captions with \underline{R}el\underline{I}able \underline{FI}n\underline{E}-grained statics and \underline{D}ynamics. Specifically, we resort to large language models (LLM) and large multimodal models (LMM) with our proposed Statics and Dynamics Enhanced Captioning modules to generate diverse fine-grained captions for each video. To filter out the inaccurate annotations caused by the LLM hallucination, we propose a Fine-Granularity Aware Noise Evaluator where we fine-tune a video foundation model with disturbed hard-negatives augmented contrastive and matching losses. With VERIFIED, we construct a more challenging fine-grained VCMR benchmark containing Charades-FIG, DiDeMo-FIG, and ActivityNet-FIG which demonstrate a high level of annotation quality. We evaluate several state-of-the-art VCMR models on the proposed dataset, revealing that there is still significant scope for fine-grained video understanding in VCMR. Code and Datasets are in \href{https://github.com/hlchen23/VERIFIED}{https://github.com/hlchen23/VERIFIED}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。