arXiv:2410.14265cs.CV2024-10被引 1

让静物生成更精准:专注前景一致性,避免背景干扰。

HYPNOS : Highly Precise Foreground-focused Diffusion Finetuning for Inanimate Objects

  • 用内容导向提示词+前景判别模块,分离前景与背景
  • 在少量样本下实现高保真静物生成,背景多样性不降低
  • 适合需要精细控制物体外观的个性化图像生成场景

近年来,基于扩散模型的个性化文生图任务成为计算机视觉研究热点。一个稳健的扩散模型应能在少量相关输入样本下,实现对特定产品结果的近乎完美重建。然而,当前主流的扩散微调技术在保持前景对象一致性方面表现不足,且受限于生成多样背景,导致过拟合问题:输入提示信息可能被模糊传递至前景和背景区域,而非仅作用于背景。为此,我们提出 Hypnos,一种高精度的前景聚焦型扩散微调方法。该方法针对静物生成任务,采用两种核心策略:(i) 内容导向提示词设计,(ii) 引入额外的前景聚焦判别模块,并通过所提出的监督机制与扩散模型联合微调。二者结合使模型具备前景-背景解耦能力。实验表明,相比原有方法,Hypnos 在鲁棒性与视觉效果上均有显著提升。我们还开展了多维度分析,深入揭示个性化训练在不同条件下的行为特征。

原文摘要 · Abstract (English)

In recent years, personalized diffusion-based text-to-image generative tasks have been a hot topic in computer vision studies. A robust diffusion model is determined by its ability to perform near-perfect reconstruction of certain product outcomes given few related input samples. Unfortunately, the current prominent diffusion-based finetuning technique falls short in maintaining the foreground object consistency while being constrained to produce diverse backgrounds in the image outcome. In the worst scenario, the overfitting issue may occur, meaning that the foreground object is less controllable due to the condition above, for example, the input prompt information is transferred ambiguously to both foreground and background regions, instead of the supposed background region only. To tackle the issues above, we proposed Hypnos, a highly precise foreground-focused diffusion finetuning technique. On the image level, this strategy works best for inanimate object generation tasks, and to do so, Hypnos implements two main approaches, namely: (i) a content-centric prompting strategy and (ii) the utilization of our additional foreground-focused discriminative module. The utilized module is connected with the diffusion model and finetuned with our proposed set of supervision mechanism. Combining the strategies above yielded to the foreground-background disentanglement capability of the diffusion model. Our experimental results showed that the proposed strategy gave a more robust performance and visually pleasing results compared to the former technique. For better elaborations, we also provided extensive studies to assess the fruitful outcomes above, which reveal how personalization behaves in regard to several training conditions.

扩散模型静物生成前景聚焦微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。