通过空间编织注意力增强文生图模型,实现高质量梗图视频生成。
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
- 引入空间编织注意力机制优化2D特征图的注意力结构
- 在梗图视频生成任务中显著提升图像质量与语义一致性
- 兼容SD1.5衍生模型,适合开源社区复现与二次开发
我们提出一种向文本到图像基础模型插入适配器的有效方法,可在保持基座模型泛化能力的同时执行复杂下游任务。该方法的核心思想是优化与二维特征图相关的注意力机制,从而提升适配器性能。该方法在梗图视频生成任务上得到验证,取得了显著成果。我们希望这项工作能为大型文生图模型的后训练任务提供新思路。此外,由于该方法对SD1.5衍生模型具有良好兼容性,对开源社区具有实际价值。因此,我们将公开相关代码(https://songkey.github.io/hellomeme)。
原文摘要 · Abstract (English)
We propose an effective method for inserting adapters into text-to-image foundation models, which enables the execution of complex downstream tasks while preserving the generalization ability of the base model. The core idea of this method is to optimize the attention mechanism related to 2D feature maps, which enhances the performance of the adapter. This approach was validated on the task of meme video generation and achieved significant results. We hope this work can provide insights for post-training tasks of large text-to-image models. Additionally, as this method demonstrates good compatibility with SD1.5 derivative models, it holds certain value for the open-source community. Therefore, we will release the related code (\url{https://songkey.github.io/hellomeme}).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。