arXiv:2411.15580cs.CV2024-11CVPR被引 10

无需微调即可生成指定背景色的前景图像,实现自由分离前景与背景。

TKG-DM: Training-free Chroma Key Content Generation Diffusion Model

  • 通过优化初始噪声的颜色成分,直接控制生成图像的背景色。
  • 在定性和定量评估中超越现有方法,媲美甚至超过微调模型效果。
  • 适用于文本到视频、一致性模型等任务,适合需要独立控制前景背景的场景。

扩散模型已能生成高保真、文本忠实度高的图像,但以 Stable Diffusion 为代表的大型文生图模型难以生成前景物体置于特定色背景上的图像,导致无法在不微调的情况下分离前景与背景。为此,我们提出首个无需训练的染色键内容生成扩散模型(TKG-DM),通过优化初始随机噪声,生成前景物体位于可指定颜色背景上的图像。该方法首次探索了在初始噪声中操控颜色特征以实现可控背景生成,无需微调即可实现前景与背景的精确分离。大量实验表明,该训练免费方法在定性与定量评价上均优于现有方法,达到甚至超越微调模型水平。此外,我们成功将其扩展至一致性模型和文生视频任务,凸显其在需独立控制前景与背景的生成应用中的变革潜力。

原文摘要 · Abstract (English)

Diffusion models have enabled the generation of high-quality images with a strong focus on realism and textual fidelity. Yet, large-scale text-to-image models, such as Stable Diffusion, struggle to generate images where foreground objects are placed over a chroma key background, limiting their ability to separate foreground and background elements without fine-tuning. To address this limitation, we present a novel Training-Free Chroma Key Content Generation Diffusion Model (TKG-DM), which optimizes the initial random noise to produce images with foreground objects on a specifiable color background. Our proposed method is the first to explore the manipulation of the color aspects in initial noise for controlled background generation, enabling precise separation of foreground and background without fine-tuning. Extensive experiments demonstrate that our training-free method outperforms existing methods in both qualitative and quantitative evaluations, matching or surpassing fine-tuned models. Finally, we successfully extend it to other tasks (e.g., consistency models and text-to-video), highlighting its transformative potential across various generative applications where independent control of foreground and background is crucial.

扩散模型图像生成前景分离无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。