为文生视频优化视频描述,提升生成质量
VC4VG: Optimizing Video Captions for Text-to-Video Generation
- 从视频生成需求出发,拆解描述关键要素并设计优化方法
- 构建多维度评估基准,发现描述质量与生成效果强相关
- 适合文生视频模型训练者使用,助力提升生成一致性
近期文生视频(T2V)技术的发展凸显了高质量视频-文本对在训练生成连贯、指令对齐视频模型中的关键作用。然而,针对T2V训练专门优化视频描述的策略仍不充分。本文提出VC4VG(Video Captioning for Video Generation),一个面向T2V模型需求的完整描述优化框架。我们从T2V视角分析描述内容,将视频重建所需的核心要素分解为多个维度,并提出系统化的描述设计方法。为支持评估,构建了VC4VG-Bench,包含细粒度、多维度且按必要性分级的指标,贴合T2V特定要求。大量T2V微调实验表明,描述质量提升与视频生成性能显著正相关,验证了该方法的有效性。所有基准工具与代码已开源:https://github.com/alimama-creative/VC4VG。
原文摘要 · Abstract (English)
Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video captions specifically for T2V training remain underexplored. In this paper, we introduce VC4VG (Video Captioning for Video Generation), a comprehensive caption optimization framework tailored to the needs of T2V models. We begin by analyzing caption content from a T2V perspective, decomposing the essential elements required for video reconstruction into multiple dimensions, and proposing a principled caption design methodology. To support evaluation, we construct VC4VG-Bench, a new benchmark featuring fine-grained, multi-dimensional, and necessity-graded metrics aligned with T2V-specific requirements. Extensive T2V fine-tuning experiments demonstrate a strong correlation between improved caption quality and video generation performance, validating the effectiveness of our approach. We release all benchmark tools and code at https://github.com/alimama-creative/VC4VG to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。