探索类SORA视频生成模型的能力与局限,推动高质量视频创作技术发展
The Dawn of Video Generation: Preliminary Explorations with SORA-like Models
- 采用扩散模型架构,实现高分辨率、自然运动的长视频生成
- 在文本到视频生成中展现更强视觉-语言对齐与可控性
- 揭示当前评测体系难以匹配人类偏好,需改进评估方法
高质量视频生成(包括文本到视频、图像到视频、视频到视频)在内容创作和世界模拟中具有重要意义,使人们能以新方式表达创造力并理解世界。类SORA模型通过从UNet向更可扩展、参数量更大的DiT架构演进,结合大规模数据与优化训练策略,在高分辨率、自然运动、视觉-语言对齐及长视频可控性方面取得显著进步。然而,尽管已有闭源与开源的DiT模型涌现,对其能力与局限的系统性研究仍不足。同时,现有基准难以全面覆盖此类模型进展,评估指标也常与人类偏好不一致。
原文摘要 · Abstract (English)
High-quality video generation, encompassing text-to-video (T2V), image-to-video (I2V), and video-to-video (V2V) generation, holds considerable significance in content creation to benefit anyone express their inherent creativity in new ways and world simulation to modeling and understanding the world. Models like SORA have advanced generating videos with higher resolution, more natural motion, better vision-language alignment, and increased controllability, particularly for long video sequences. These improvements have been driven by the evolution of model architectures, shifting from UNet to more scalable and parameter-rich DiT models, along with large-scale data expansion and refined training strategies. However, despite the emergence of DiT-based closed-source and open-source models, a comprehensive investigation into their capabilities and limitations remains lacking. Furthermore, the rapid development has made it challenging for recent benchmarks to fully cover SORA-like models and recognize their significant advancements. Additionally, evaluation metrics often fail to align with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。