通过分阶段采样融合多概念,提升个性化图像视频生成质量
TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation
- 分两阶段采样:前期保目标对象,后期用Tweedie公式融合外观
- 生成多概念图像视频的保真度显著高于现有方法
- 可无缝扩展至图像到视频生成,适合需要多角色定制的场景
尽管文本到图像和视频生成模型在个性化定制方面取得显著进展,但有效融合多个个性化概念仍具挑战。为此,我们提出TweedieMix,一种在推理阶段组合定制化扩散模型的新方法。通过分析逆向扩散采样的特性,该方法将采样过程分为两个阶段:初始阶段采用多对象感知采样技术,确保目标对象被包含;后期在去噪图像空间中利用Tweedie公式融合自定义概念的外观。实验表明,TweedieMix生成的多概念图像视频保真度显著优于现有方法。此外,该框架可轻松扩展至图像到视频扩散模型,实现包含多个个性化概念的视频生成。结果与源代码见匿名项目页面。
原文摘要 · Abstract (English)
Despite significant advancements in customizing text-to-image and video generation models, generating images and videos that effectively integrate multiple personalized concepts remains a challenging task. To address this, we present TweedieMix, a novel method for composing customized diffusion models during the inference phase. By analyzing the properties of reverse diffusion sampling, our approach divides the sampling process into two stages. During the initial steps, we apply a multiple object-aware sampling technique to ensure the inclusion of the desired target objects. In the later steps, we blend the appearances of the custom concepts in the de-noised image space using Tweedie's formula. Our results demonstrate that TweedieMix can generate multiple personalized concepts with higher fidelity than existing methods. Moreover, our framework can be effortlessly extended to image-to-video diffusion models, enabling the generation of videos that feature multiple personalized concepts. Results and source code are in our anonymous project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。