arXiv:2410.13370cs.CVcs.AI2024-10IJCAI被引 19

让文字生成图像可精准调整局部元素,提升个性化创作能力。

MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

  • 通过动态掩码降解抑制无关视觉信息干扰
  • 采用双流平衡机制解决组件学习不均问题
  • 适合需要精细控制图像局部的创意设计用户

文本到图像扩散模型虽能生成高质量图像,但对视觉概念的细粒度控制能力不足,限制了创造力。为此,我们提出组件可控个性化这一新任务,使用户能够自定义并重新配置概念中的单个组件。该任务面临两大挑战:语义污染(无关元素干扰目标概念)与语义不平衡(目标概念与组件学习比例失衡)。为此,我们设计 MagicTailor 框架,采用动态掩码降解方法自适应地扰动不需要的视觉语义,并引入双流平衡机制实现更均衡的期望视觉语义学习。实验结果表明,MagicTailor 在该任务上表现优异,支持更具个性化和创造性的图像生成。

原文摘要 · Abstract (English)

Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. This task faces two challenges: semantic pollution, where undesired elements disrupt the target concept, and semantic imbalance, which causes disproportionate learning of the target concept and component. To address these, we design MagicTailor, a framework that uses Dynamic Masked Degradation to adaptively perturb unwanted visual semantics and Dual-Stream Balancing for more balanced learning of desired visual semantics. The experimental results show that MagicTailor achieves superior performance in this task and enables more personalized and creative image generation.

图像生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。