arXiv:2607.18789cs.CV2026-07

通过可控数据集研究文本到视频生成中训练数据质量的影响。

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

论文配图:Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
图 1 · 摘自论文原文
  • 构建可编程测试平台Moving Alphabet,精确控制视频内容与标注质量。
  • 数据分布多样性和标注准确性显著影响模型泛化与训练效率。
  • 高质量数据对预训练至关重要,微调无法完全弥补低质数据缺陷。

过去五年间,文本到视频生成得益于模型规模、数据量和计算资源的扩展而取得显著进展。与模型架构不同,训练数据常被忽视。真实世界的数据筛选复杂,涉及从原始视频中选取片段并进行标注,以构建用于学习文本到视频映射的视频-文本对。本文研究数据分布和标注质量对文本到视频模型的影响。为实现受控实验,我们引入Moving Alphabet——一个程序化测试平台,通过渲染不同字体、颜色、大小和位置的字母,在黑色背景上以不同方向和速度移动,从而精确控制数据分布和标注质量(通过破坏真实元数据实现)。实验得出三个发现:a)视频内容和时长的多样化与均衡分布对泛化能力至关重要;b)标注质量显著影响模型性能与训练效率,表明文本到视频模型受限于视频理解能力;c)无分类器引导和在高质量数据上的微调可部分恢复因标注损坏导致的性能下降,但无法完全补偿预训练数据质量差的问题。我们认为这些见解有助于大型文本到视频模型的开发,并呼吁加强对预训练数据科学的关注。

原文摘要 · Abstract (English)

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.

文本生成视频训练数据可控实验数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。