系统梳理扩散模型生成视频的技术框架与核心挑战
Image-to-Video Diffusion: From Foundations to Open Frontiers

- 按架构与训练范式分类现有图像转视频方法
- 提炼出条件编码、时序建模等四大核心设计要素
- 适合想深入视频生成研究的学者与工程师参考
基于扩散模型的图像到视频(I2V)生成已成为生成模型的重要方向,能够将参考图像(可选附加条件)转化为时间连贯的视频。相比更广泛的视频生成任务,该任务对内容一致性、身份保留和运动连贯性要求更高。尽管相关研究迅速发展,但多数工作仍将其置于更广泛的主题下讨论,缺乏针对 I2V 的专门分类体系与系统性分析。本文首次将扩散 I2V 生成作为独立课题进行研究,系统回顾任务定义、模型架构、数据集与评估指标,并基于架构与训练范式构建分类体系。进一步提炼出四大核心设计:条件编码、时序建模、噪声先验设计与时空上采样,并探讨代表性应用场景及主要开放挑战。
原文摘要 · Abstract (English)
Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation settings, this task places stricter demands on content consistency, identity preservation, and motion coherence. Although the literature grows rapidly, existing works mostly discuss I2V generation within broader topics and still lack a dedicated taxonomy together with a systematic analysis centered on this field. This work addresses that gap by treating diffusion I2V generation as a standalone subject. It first reviews the task formulation, model architectures, datasets, and evaluation metrics, and then organizes existing methods through a taxonomy based on architecture and training paradigm. It further distills four core designs, namely condition encoding, temporal modeling, noise prior design, and spatial-temporal upsampling, and discusses representative application scenarios together with major open challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。