为可控文生视频设计了全面的视频描述评估基准,提升模型训练效果。
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
- 构建多维度标注体系,涵盖美学、内容、运动与物理规律
- 分自动与人工评估模块,兼顾效率与准确性
- 验证显示评分与文生视频质量高度相关,适合模型优化
可控文生视频(T2V)模型的训练依赖于视频与描述之间的对齐,但现有研究极少将视频描述评估与T2V生成评估关联。本文提出VidCapBench,一个专为T2V生成设计的视频描述评估框架,不依赖特定描述格式。通过结合专家模型标注与人工精修的数据标注流程,为每段视频标注涵盖视频美学、内容、运动及物理规律的关键信息。VidCapBench将这些属性划分为可自动评估与需人工评估两部分,满足敏捷开发中的快速评估与严谨验证的需求。对多个主流描述生成模型的评估表明,VidCapBench在稳定性和全面性上优于现有方法。使用现成T2V模型的验证显示,其得分与T2V质量指标存在显著正相关,证明VidCapBench能有效指导T2V模型训练。项目已开源:https://github.com/VidCapBench/VidCapBench。
原文摘要 · Abstract (English)
The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces VidCapBench, a video caption evaluation scheme specifically designed for T2V generation, agnostic to any particular caption format. VidCapBench employs a data annotation pipeline, combining expert model labeling and human refinement, to associate each collected video with key information spanning video aesthetics, content, motion, and physical laws. VidCapBench then partitions these key information attributes into automatically assessable and manually assessable subsets, catering to both the rapid evaluation needs of agile development and the accuracy requirements of thorough validation. By evaluating numerous state-of-the-art captioning models, we demonstrate the superior stability and comprehensiveness of VidCapBench compared to existing video captioning evaluation approaches. Verification with off-the-shelf T2V models reveals a significant positive correlation between scores on VidCapBench and the T2V quality evaluation metrics, indicating that VidCapBench can provide valuable guidance for training T2V models. The project is available at https://github.com/VidCapBench/VidCapBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。