arXiv:2511.17615cs.CVcs.AI2025-11

无需微调即可高保真融合多个个性化概念生成图像

Plug-and-Play Multi-Concept Adaptive Blending for High-Fidelity Text-to-Image Synthesis

  • 用引导式外观注意力精准保留每个概念的外观特征
  • 通过掩码引导噪声混合,保护非个性化区域完整性
  • 提出背景稀释++策略,有效防止概念泄露

将多个个性化概念融入单张图像已成为文本到图像生成的重要方向。然而,现有方法在复杂多对象场景中常因对个性化与非个性化区域的意外修改而表现不佳,不仅破坏提示结构,还导致区域间语义不一致。为此,我们提出无需微调的即插即用多概念自适应融合方法(PnP-MIX),通过引导式外观注意力忠实呈现每个个性化概念的预期外观。为进一步提升组合保真度,引入掩码引导噪声混合策略,保护背景或无关物体等非个性化区域的同时,实现个性化物体的精确融合。最后,针对概念泄露问题,提出背景稀释++策略,显著减少个性化特征向其他区域扩散,促进特征在个性化区域内的准确定位。大量实验表明,PnP-MIX在单概念与多概念个性化场景中均优于现有方法,展现出强鲁棒性与卓越性能。

原文摘要 · Abstract (English)

Integrating multiple personalized concepts into a single image has recently become a significant area of focus within Text-to-Image (T2I) generation. However, existing methods often underperform on complex multi-object scenes due to unintended alterations in both personalized and non-personalized regions. This not only fails to preserve the intended prompt structure but also disrupts interactions among regions, leading to semantic inconsistencies. To address this limitation, we introduce plug-and-play multi-concept adaptive blending for high-fidelity text-to-image synthesis (PnP-MIX), an innovative, tuning-free approach designed to seamlessly embed multiple personalized concepts into a single generated image. Our method leverages guided appearance attention to faithfully reflect the intended appearance of each personalized concept. To further enhance compositional fidelity, we present a mask-guided noise mixing strategy that preserves the integrity of non-personalized regions such as the background or unrelated objects while enabling the precise integration of personalized objects. Finally, to mitigate concept leakage, i.e., the inadvertent leakage of personalized concept features into other regions, we propose background dilution++, a novel strategy that effectively reduces such leakage and promotes accurate localization of features within personalized regions. Extensive experimental results demonstrate that PnP-MIX consistently surpasses existing methodologies in both single- and multi-concept personalization scenarios, underscoring its robustness and superior performance without additional model tuning.

文本生成图像个性化融合无微调概念保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。