首个大规模视频字幕细粒度评估基准,助力生成更精准的图文视频。
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
- 构建涵盖21个细粒度维度的问答对,覆盖时空细节
- 包含5677个视频与10.9万+问答对,支持高精度评估
- 基于大模型自动分析,适合优化视频生成模型
视频字幕在文本到视频生成任务中至关重要,其质量直接影响生成视频的语义连贯性与视觉保真度。尽管大型视觉语言模型(VLM)在字幕生成方面展现出巨大潜力,但现有评估基准未能充分解决细粒度评价问题,尤其在捕捉对视频生成至关重要的时空细节方面存在不足。为此,我们提出细粒度视频字幕评估基准VCapsBench,这是首个大规模细粒度基准,包含5,677(5K+)个视频和109,796(100K+)个问答对。这些问答对在21个细粒度维度(如摄像机运动、镜头类型等)上系统标注,实证证明这些维度对文本到视频生成极为关键。我们进一步引入三种评估指标(准确率AR、不一致率IR、覆盖率CR),并设计基于大语言模型(LLM)的自动化评估流水线,通过对比问答对分析验证字幕质量。该基准为字幕优化提供可操作洞察,推动鲁棒性文本到视频模型的发展。数据集与代码已开源:https://github.com/GXYM/VCapsBench。
原文摘要 · Abstract (English)
Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmarks inadequately address fine-grained evaluation, particularly in capturing spatial-temporal details critical for video generation. To address this gap, we introduce the Fine-grained Video Caption Evaluation Benchmark (VCapsBench), the first large-scale fine-grained benchmark comprising 5,677 (5K+) videos and 109,796 (100K+) question-answer pairs. These QA-pairs are systematically annotated across 21 fine-grained dimensions (e.g., camera movement, and shot type) that are empirically proven critical for text-to-video generation. We further introduce three metrics (Accuracy (AR), Inconsistency Rate (IR), Coverage Rate (CR)), and an automated evaluation pipeline leveraging large language model (LLM) to verify caption quality via contrastive QA-pairs analysis. By providing actionable insights for caption optimization, our benchmark can advance the development of robust text-to-video models. The dataset and codes are available at website: https://github.com/GXYM/VCapsBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。