首个融合图文视频的统一视频超分框架,支持多模态条件生成。
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- 构建统一框架,融合文本、图像、视频三类条件输入。
- 在多个数据集上实现超越现有方法的细节还原与条件符合度。
- 可与基础模型结合生成4K级多模态引导视频,突破此前技术瓶颈。
级联视频超分辨率已成为缓解大模型生成高清视频计算负担的有前景技术。然而,现有研究主要局限于文生视频任务,未能充分利用除文本外的其他生成条件,而这些条件对多模态视频生成的保真度至关重要。为此,我们提出UniMMVSR,首个统一的生成式视频超分辨率框架,支持文本、图像和视频三类混合模态条件。我们在潜空间视频扩散模型中系统探索了条件注入策略、训练方案与数据混合技术。关键挑战在于设计差异化的数据构建与条件利用方法,以应对不同条件与目标视频间相关性的差异。实验表明,UniMMVSR显著优于现有方法,在细节表现与多模态条件一致性方面均有提升。我们还验证了将UniMMVSR与基座模型结合,实现4K级多模态引导视频生成的可行性,这是此前技术无法达成的目标。
原文摘要 · Abstract (English)
Cascaded video super-resolution has emerged as a promising technique for decoupling the computational burden associated with generating high-resolution videos using large foundation models. Existing studies, however, are largely confined to text-to-video tasks and fail to leverage additional generative conditions beyond text, which are crucial for ensuring fidelity in multi-modal video generation. We address this limitation by presenting UniMMVSR, the first unified generative video super-resolution framework to incorporate hybrid-modal conditions, including text, images, and videos. We conduct a comprehensive exploration of condition injection strategies, training schemes, and data mixture techniques within a latent video diffusion model. A key challenge was designing distinct data construction and condition utilization methods to enable the model to precisely utilize all condition types, given their varied correlations with the target video. Our experiments demonstrate that UniMMVSR significantly outperforms existing methods, producing videos with superior detail and a higher degree of conformity to multi-modal conditions. We also validate the feasibility of combining UniMMVSR with a base model to achieve multi-modal guided generation of 4K video, a feat previously unattainable with existing techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。