arXiv:2502.15203cs.CVcs.AI2025-02中稿 · IEEE SMC 2025被引 2

无需微调即可实现多概念个性化图像生成,避免干扰与失真。

FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation

  • 通过引导外观注意力增强个性化区域视觉质量。
  • 利用掩码引导噪声混合保护非个性化区域不被破坏。
  • 采用背景稀释减少概念泄露,适合实际应用需求。

将多个个性化概念融入单张图像在文本到图像(T2I)生成中受到关注。然而,现有方法常因非个性化区域的失真和需要额外微调而导致复杂场景下性能下降,限制了实用性。为此,我们提出FlipConcept,一种无需额外调优即可无缝融合多个个性化概念的新方法。引入引导外观注意力以提升个性化概念的视觉保真度;提出掩码引导噪声混合,保护非个性化区域在概念融合过程中的完整性;最后应用背景稀释策略,最小化概念泄露——即个性化概念与图像中其他物体的意外混合。实验表明,尽管无需调优,该方法在单个和多个个性化概念推理上均优于现有模型,验证了其在可扩展、高质量多概念个性化方面的有效性与实用性。

原文摘要 · Abstract (English)

Integrating multiple personalized concepts into a single image has recently gained attention in text-to-image (T2I) generation. However, existing methods often suffer from performance degradation in complex scenes due to distortions in non-personalized regions and the need for additional fine-tuning, limiting their practicality. To address this issue, we propose FlipConcept, a novel approach that seamlessly integrates multiple personalized concepts into a single image without requiring additional tuning. We introduce guided appearance attention to enhance the visual fidelity of personalized concepts. Additionally, we introduce mask-guided noise mixing to protect non-personalized regions during concept integration. Lastly, we apply background dilution to minimize concept leakage, i.e., the undesired blending of personalized concepts with other objects in the image. In our experiments, we demonstrate that the proposed method, despite not requiring tuning, outperforms existing models in both single and multiple personalized concept inference. These results demonstrate the effectiveness and practicality of our approach for scalable, high-quality multi-concept personalization.

图像生成个性化无微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。