构建细粒度视频理解评测集,揭示模型在时空感知上的偏差。
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
- 设计包含1000对视频与详细标注的评测集,支持时空分离标注。
- 提出针对检索与描述任务的新评估指标,发现模型存在时空偏差。
- 基于多模态语言模型实现统一框架,在两项任务上表现均具竞争力。
视频理解(包括视频描述与检索)仍是视频-语言模型(VLMs)的重大挑战。现有评测集仅提供简短描述,限制了对细节理解能力的评估。为此,我们提出CaReBench,一个面向细粒度视频描述与检索的评测基准,包含1000组高质量视频与人工标注的详细描述。其独特之处在于为每段视频提供手动分离的空间与时间标注。基于此设计,我们引入两个专用评估指标:ReBias(用于检索)与CapST(用于描述),可全面分析VLM中固有的时空偏差。此外,为统一处理检索与描述任务,我们基于多模态语言模型(MLLM)构建简单基线,通过两阶段监督微调(SFT)充分释放其潜力,使其不仅能生成详细视频描述,还能提取视频特征。实验表明,相比专为检索设计的CLIP模型和擅长描述的主流MLLM,该基线在细粒度检索与详细描述任务上均表现出色。
原文摘要 · Abstract (English)
Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, we present CaReBench, a testing benchmark for fine-grained video captioning and retrieval with 1,000 high-quality pairs of videos and human-annotated detailed captions. Uniquely, it provides manually separated spatial annotations and temporal annotations for each video. Based on this design, we introduce two evaluation metrics, ReBias and CapST, specifically tailored for video retrieval and video captioning tasks, respectively. These metrics enable a comprehensive investigation into the spatial and temporal biases inherent in VLMs. In addition, to handle both video retrieval and video captioning tasks in a unified framework, we develop a simple baseline based on a Multimodal Language Model (MLLM). By implementing a two-stage Supervised Fine-Tuning (SFT), we fully unlock the potential of MLLM, enabling it not only to generate detailed video descriptions but also to extract video features. Surprisingly, experimental results demonstrate that, compared to the CLIP-based models designed for retrieval and the popular MLLMs skilled in video captioning, our baseline shows competitive performance in both fine-grained video retrieval and video detailed captioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。