arXiv:2507.01938cs.CV2025-07被引 3

构建首个连贯多片段图文视频数据集,支持长序列故事生成。

CI-VID: A Coherent Interleaved Text-Video Dataset

  • 设计连贯交织的文本-视频对,包含跨片段语义衔接
  • 超34万样本,显著提升模型生成视频的时序一致性
  • 适配故事化视频生成任务,适合研究长序列生成的学者

文本到视频(T2V)生成近年来受到广泛关注,推动了大量高质量数据集的发展。然而,现有公开数据集主要由孤立的文本-视频对构成,难以支持连贯多片段视频序列建模。为解决这一问题,我们提出CI-VID数据集,从孤立T2V生成迈向文本与视频到视频(TV2V)生成,使模型能够生成具有连贯性、多场景的视频序列。CI-VID包含超过34万条样本,每条样本均包含一组语义连贯的视频片段,并配有描述各片段内容及片段间过渡关系的文本标题,实现视觉与文本的双重引导。为验证其有效性,我们设计了涵盖人工评估、基于视觉语言模型的评估和相似性度量的多维度基准。实验表明,基于CI-VID训练的模型在生成视频序列时,在准确性和内容一致性方面均有显著提升,有助于生成具有流畅视觉过渡和强时间连贯性的叙事性内容,凸显了该数据集的质量与实用性。我们已公开CI-VID数据集及数据构建与评估代码:https://github.com/ymju-BAAI/CI-VID

原文摘要 · Abstract (English)

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed of isolated text-video (T-V) pairs and thus fail to support the modeling of coherent multi-clip video sequences. To address this limitation, we introduce CI-VID, a dataset that moves beyond isolated text-to-video (T2V) generation toward text-and-video-to-video (TV2V) generation, enabling models to produce coherent, multi-scene video sequences. CI-VID contains over 340,000 samples, each featuring a coherent sequence of video clips with text captions that capture both the individual content of each clip and the transitions between them, enabling visually and textually grounded generation. To further validate the effectiveness of CI-VID, we design a comprehensive, multi-dimensional benchmark incorporating human evaluation, VLM-based assessment, and similarity-based metrics. Experimental results demonstrate that models trained on CI-VID exhibit significant improvements in both accuracy and content consistency when generating video sequences. This facilitates the creation of story-driven content with smooth visual transitions and strong temporal coherence, underscoring the quality and practical utility of the CI-VID dataset We release the CI-VID dataset and the accompanying code for data construction and evaluation at: https://github.com/ymju-BAAI/CI-VID

视频生成多片段数据集叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。