评测视频模型生成长段落描述的能力,填补现有短片段评测空白。
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

- 基于5小时电影片段构建90秒长视频的段落级描述评估集
- 用5个先进LLM嵌入模型融合评分,提升评测可靠性
- 适合研究长文本生成、视频理解与多模态模型评估的学者
现有视频-语言模型评测主要聚焦于短片段和单句指标,未能检验系统生成准确长篇段落描述的能力。本文提出CLIP-CC-Bench,一个基于5小时电影内容构建的长篇视频描述评估套件,将内容分割为90秒片段,每个片段配以专家撰写的段落式参考描述。评估采用五种最先进的基于LLM的嵌入模型集成,增强可靠性并减少单一模型偏差,并结合粗粒度与细粒度语义匹配两种互补方法,对比模型生成描述与参考描述。在此框架下,评估了17个顶尖视频-语言模型,报告其Borda聚合排名与平均得分。通过评判者间一致性与自助法排名稳定性量化协议内部可靠性。项目代码、模型输出及聚合工具已开源,支持可复现性。CLIP-CC-Bench为长篇视频描述提供了实用评估框架,弥补了现有短片段与问答类评测的空白。
原文摘要 · Abstract (English)
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。