arXiv:2411.19261cs.CV2024-11被引 4

解决多主体图像生成中主体混淆与位置错位问题

Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention

  • 通过隔离注意力防止多个主体相互干扰
  • 通过重定位使参考与目标主体位置对齐
  • 无需训练,显著提升开放域多主体一致性

无需训练的扩散模型在开放域多主体图像生成中已取得显著进展,其核心思想是在注意力层中引入参考主体信息。然而,现有方法在处理多个主体时仍表现不佳。本文揭示两个关键问题:一是目标图像中不同主体间的非期望吸引力导致多个主体融合为单一实体;二是标记倾向于参考邻近标记,当参考与目标图像中主体位置差异较大时,注意力机制效果下降。为此,提出无需训练的扩散模型IR-Diffusion,采用隔离注意力确保目标图像中多个主体互不引用,有效避免主体融合;同时通过缩放与重定位,使参考与目标图像中的主体处于相同位置,从而增强一致性。大量实验表明,IR-Diffusion在开放域场景下显著提升多主体一致性,优于所有现有方法。

原文摘要 · Abstract (English)

Training-free diffusion models have achieved remarkable progress in generating multi-subject consistent images within open-domain scenarios. The key idea of these methods is to incorporate reference subject information within the attention layer. However, existing methods still obtain suboptimal performance when handling numerous subjects. This paper reveals two primary issues contributing to this deficiency. Firstly, the undesired internal attraction between different subjects within the target image can lead to the convergence of multiple subjects into a single entity. Secondly, tokens tend to reference nearby tokens, which reduces the effectiveness of the attention mechanism when there is a significant positional difference between subjects in reference and target images. To address these issues, we propose a training-free diffusion model with Isolation and Reposition Attention, named IR-Diffusion. Specifically, Isolation Attention ensures that multiple subjects in the target image do not reference each other, effectively eliminating the subject convergence. On the other hand, Reposition Attention involves scaling and repositioning subjects in both reference and target images to the same position within the images. This ensures that subjects in the target image can better reference those in the reference image, thereby maintaining better consistency. Extensive experiments demonstrate that IR-Diffusion significantly enhances multi-subject consistency, outperforming all existing methods in open-domain scenarios.

扩散模型图像生成多主体一致注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。