构建首个长视频细粒度检索基准,解决标注质量与跨度难题
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
- 设计多阶段生成框架,提升长视频细粒度字幕质量与可扩展性
- 包含467段长视频、超4万条片段,平均时长显著超过现有数据集
- 适合研究长视频理解、跨模态检索与高质量标注方法的学者
长视频蕴含海量信息,使其在多模态学习中成为视频-文本检索的关键挑战。现有基准普遍存在视频时长短、字幕质量差、标注粒度粗等问题,制约了先进方法的评估。为此,我们提出LoVR,一个专为长视频文本检索设计的基准。该数据集包含467段长视频和超过40,804条细粒度片段,配有高质量字幕。为克服机器生成标注质量低的问题,我们提出一种高效字幕生成框架,融合视觉语言模型(VLM)自动生成、字幕质量评分与动态优化机制,兼顾准确性与可扩展性。此外,引入语义融合方法,生成连贯且不丢失关键上下文的全视频字幕。LoVR在视频长度、标注精细度与数据规模上均显著超越现有基准,为视频理解与检索带来新挑战。对多种先进嵌入模型的实验表明,该基准极具挑战性,暴露出当前方法的局限,为未来研究提供重要参考。代码与数据集已公开:https://lovrbench.github.io/
原文摘要 · Abstract (English)
Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality captions, and coarse annotation granularity, which hinder the evaluation of advanced video-text retrieval methods. To address these limitations, we introduce LoVR, a benchmark specifically designed for long video-text retrieval. LoVR contains 467 long videos and over 40,804 fine-grained clips with high-quality captions. To overcome the issue of poor machine-generated annotations, we propose an efficient caption generation framework that integrates VLM automatic generation, caption quality scoring, and dynamic refinement. This pipeline improves annotation accuracy while maintaining scalability. Furthermore, we introduce a semantic fusion method to generate coherent full-video captions without losing important contextual information. Our benchmark introduces longer videos, more detailed captions, and a larger-scale dataset, presenting new challenges for video understanding and retrieval. Extensive experiments on various advanced embedding models demonstrate that LoVR is a challenging benchmark, revealing the limitations of current approaches and providing valuable insights for future research. We release the code and dataset link at https://lovrbench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。