系统梳理文本生成视频技术演进与挑战,助力高质量可控视频生成。
Bridging Text and Video Generation: A Survey
- 从对抗生成到扩散-变换器架构,揭示模型演进逻辑
- 总结主流数据集与训练配置,支持复现与资源规划
- 提出评价体系不足与未来研究方向,适合研究者参考
文本到视频(T2V)生成技术有望在教育、营销、娱乐及视觉/阅读障碍辅助等领域带来变革,通过自然语言提示生成连贯视觉内容。该领域从早期的GAN和VAE发展至基于扩散的混合扩散-变换器(DiT)架构,显著提升了输出质量与时间一致性。然而仍面临语义对齐、长程连贯性与计算效率等挑战。本文全面综述了T2V生成模型的发展历程,涵盖从早期生成对抗网络到现代扩散-变换器架构的技术演进,解析其工作机制、克服的局限及其必要性。系统梳理了训练与评估所用数据集,详细列出硬件配置、GPU数量、批量大小、学习率、优化器、训练轮数等关键超参数,以支持可复现性与训练可行性分析。同时,归纳常用评估指标及其在标准基准上的表现,并讨论现有指标的局限性,指出向更符合人类感知的综合评估策略转变的趋势。最后,基于分析提出当前开放挑战与有前景的未来方向,为后续研究提供蓝图。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating coherent visual content from natural language prompts. From its inception, the field has advanced from adversarial models to diffusion-based models, yielding higher-fidelity, temporally consistent outputs. Yet challenges persist, such as alignment, long-range coherence, and computational efficiency. Addressing this evolving landscape, we present a comprehensive survey of text-to-video generative models, tracing their development from early GANs and VAEs to hybrid Diffusion-Transformer (DiT) architectures, detailing how these models work, what limitations they addressed in their predecessors, and why shifts toward new architectural paradigms were necessary to overcome challenges in quality, coherence, and control. We provide a systematic account of the datasets, which the surveyed text-to-video models were trained and evaluated on, and, to support reproducibility and assess the accessibility of training such models, we detail their training configurations, including their hardware specifications, GPU counts, batch sizes, learning rates, optimizers, epochs, and other key hyperparameters. Further, we outline the evaluation metrics commonly used for evaluating such models and present their performance across standard benchmarks, while also discussing the limitations of these metrics and the emerging shift toward more holistic, perception-aligned evaluation strategies. Finally, drawing from our analysis, we outline the current open challenges and propose a few promising future directions, laying out a perspective for future researchers to explore and build upon in advancing T2V research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。