arXiv:2412.09389cs.CVcs.AI2024-12AAAI

用统一帧组织器提升扩散模型视频生成的连贯性与画质

UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer

  • 提出非侵入式插件UFO,通过可调强度适配器增强前后景一致性
  • 在不修改原模型参数下显著改善视频画质与时间连贯性
  • 模块化设计支持多UFO组合与跨模型迁移,适合定制化视频生成

近期基于扩散的视频生成模型取得了显著进展,但普遍存在前后景一致性弱、随时间推移画质下降等问题。受美学原则启发,我们提出一种兼容任意扩散视频生成模型的非侵入式插件Uniform Frame Organizer(UFO)。UFO由一系列可调强度的自适应适配器组成,集成后无需修改原模型参数即可显著提升视频中前景与背景的一致性,并改善整体画质。UFO训练简单高效,资源消耗少,支持风格化训练。其模块化设计允许多个UFO组合使用,实现个性化视频生成模型定制。此外,UFO可在同规格不同模型间直接迁移,无需重新训练。实验表明,UFO在公开视频生成基准上有效提升生成质量,表现优于现有方法。代码将开源于https://github.com/Delong-liu-bupt/UFO。

原文摘要 · Abstract (English)

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks. The code will be publicly available at https://github.com/Delong-liu-bupt/UFO.

视频生成扩散模型一致性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。