用统一时空自回归框架实现高效高分辨率图像视频生成
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- 纯离散架构联合建模空间与时间依赖关系
- 在VBench上达83.74分,生成5秒720p视频速度超扩散模型10倍
- 首个可生成工业级720p视频的离散自回归模型,适合高效视频生成研究
我们提出InfinityStar,一种用于高分辨率图像与动态视频合成的统一时空自回归框架。基于自回归建模在视觉与语言任务中的成功,我们的纯离散方法在单一架构中联合捕捉空间与时间依赖。该统一设计自然支持文本到图像、文本到视频、图像到视频及长时交互视频生成,仅通过简单的时序自回归实现。大量实验表明,InfinityStar在VBench上得分83.74,显著优于所有自回归模型,甚至超越部分扩散模型如HunyuanVideo。无需额外优化,其生成5秒720p视频的速度约为领先扩散方法的10倍。据我们所知,InfinityStar是首个能生成工业级720p视频的离散自回归视频生成器。我们已开源全部代码与模型,以推动高效高质量视频生成研究。
原文摘要 · Abstract (English)
We introduce InfinityStar, a unified spacetime autoregressive framework for high-resolution image and dynamic video synthesis. Building on the recent success of autoregressive modeling in both vision and language, our purely discrete approach jointly captures spatial and temporal dependencies within a single architecture. This unified design naturally supports a variety of generation tasks such as text-to-image, text-to-video, image-to-video, and long interactive video synthesis via straightforward temporal autoregression. Extensive experiments demonstrate that InfinityStar scores 83.74 on VBench, outperforming all autoregressive models by large margins, even surpassing some diffusion competitors like HunyuanVideo. Without extra optimizations, our model generates a 5s, 720p video approximately 10x faster than leading diffusion-based methods. To our knowledge, InfinityStar is the first discrete autoregressive video generator capable of producing industrial level 720p videos. We release all code and models to foster further research in efficient, high-quality video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。