高效生成2K高清图像转视频,速度提升200倍且保持细节真实。
SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

- 分段条件生成:按片段逐步合成,控制每步显存开销。
- 2K视频生成仅需单张数据中心或消费级显卡,效率提升202倍。
- 结合双向上下文增强连贯性,避免细节幻觉,忠实还原输入图像。
高分辨率图像转视频(I2V)旨在生成逼真的时序动态同时保留输入图像的精细外观细节。在2K分辨率下挑战极大,现有方法存在两大缺陷:1)端到端模型内存与延迟成本过高;2)先低分辨率生成再通用超分辨率的方法易产生细节幻觉,且偏离输入特定局部结构,因超分辨率未显式依赖输入图像。为此,我们提出SwiftI2V,一种专为高分辨率I2V设计的高效框架。沿用主流两阶段设计,通过生成低分辨率运动参考以降低令牌消耗、减轻建模负担,再基于该运动引导强图像约束的2K细节重建,实现可控开销下的输入忠实恢复。为提升可扩展性,引入条件分段生成(CSG),分段合成并控制每步令牌预算,并在每段内采用双向上下文交互以增强跨段一致性与输入保真度。在VBench-I2V 2K测试中,性能媲美端到端基线,总GPU时间减少202倍。尤其支持在单张数据中心GPU(如H800)或消费级显卡(如RTX 4090)上实现实用化的2K I2V生成。
原文摘要 · Abstract (English)
High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of the input image. At 2K resolution, it becomes extremely challenging, and existing solutions suffer from various weaknesses: 1) end-to-end models are often prohibitively expensive in memory and latency; 2) cascading low-resolution generation with a generic video super-resolution tends to hallucinate details and drift from input-specific local structures, since the super-resolution stage is not explicitly conditioned on the input image. To this end, we propose SwiftI2V, an efficient framework tailored for high-resolution I2V. Following the widely used two-stage design, it addresses the efficiency--fidelity dilemma by first generating a low-resolution motion reference to reduce token costs and ease the modeling burden, then performing a strongly image-conditioned 2K synthesis guided by the motion to recover input-faithful details with controlled overhead. Specifically, to make generation more scalable, SwiftI2V introduces Conditional Segment-wise Generation (CSG) to synthesize videos segment-by-segment with a bounded per-step token budget, and adopts bidirectional contextual interaction within each segment to improve cross-segment coherence and input fidelity. On VBench-I2V at 2K resolution, SwiftI2V achieves performance comparable to end-to-end baselines while reducing total GPU-time by 202x. Particularly, it enables practical 2K I2V generation on a single datacenter GPU (e.g., H800) or consumer GPU (e.g., RTX 4090).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。