arXiv:2505.01428cs.CV2025-05CVPR被引 6

无需微调,同时用图文控制图像生成,还能避免主体混淆。

Multi-party Collaborative Attention Control for Image Customization

  • 在自注意力层中设计双操作,协调多路扩散过程。
  • 零样本定制下生成质量优于现有方法,背景一致且主体清晰。
  • 适合需要快速图文联合控制的图像生成场景。

扩散模型的快速发展推动了图像定制化需求,但现有方法存在诸多局限:通常仅支持图像或文本单条件输入;复杂视觉场景中易出现主体泄露或混淆;图像条件输出常伴随背景不一致;且计算成本高。为此,本文提出无需微调的多参与方协同注意力控制(MCA-Ctrl)方法,支持文本与复杂视觉条件共同驱动高质量图像定制。该方法在自注意力层引入两项关键操作,协调多个并行扩散过程,引导目标图像生成。通过捕捉特定主体的内容与外观,同时保持与输入条件的语义一致性。为缓解复杂场景下的主体混淆问题,进一步设计主体定位模块,基于用户指令提取精确主体及可编辑图像层。大量定量与人工评估实验表明,MCA-Ctrl在零样本图像定制任务中表现更优,有效解决前述问题。

原文摘要 · Abstract (English)

The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leakage or confusion; 3) image-conditioned outputs tend to suffer from inconsistent backgrounds; and 4) high computational costs. To address these issues, this paper introduces Multi-party Collaborative Attention Control (MCA-Ctrl), a tuning-free method that enables high-quality image customization using both text and complex visual conditions. Specifically, MCA-Ctrl leverages two key operations within the self-attention layer to coordinate multiple parallel diffusion processes and guide the target image generation. This approach allows MCA-Ctrl to capture the content and appearance of specific subjects while maintaining semantic consistency with the conditional input. Additionally, to mitigate subject leakage and confusion issues common in complex visual scenarios, we introduce a Subject Localization Module that extracts precise subject and editable image layers based on user instructions. Extensive quantitative and human evaluation experiments show that MCA-Ctrl outperforms existing methods in zero-shot image customization, effectively resolving the mentioned issues.

图像生成扩散模型多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。