arXiv:2512.23537cs.CVcs.AI2025-12被引 2

无需训练即可实现多主体布局引导生成,解决主体错乱问题

AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization

  • 自底向上解耦注意力,分离文本与图像条件的交叉注意力
  • 局部解耦确保每个主体仅关注指定区域,避免冲突与身份丢失
  • 基于预训练适配器提取特征,无需微调,支持多主体扩展

多主体定制旨在将多个用户指定主体合成到一致图像中。为解决主体缺失或冲突问题,现有方法引入布局引导以提供显式空间约束。然而,现有方法仍难以平衡文本对齐、主体身份保持和布局控制三项目标,且依赖额外训练限制了可扩展性与效率。本文提出AnyMS,一种面向布局引导的无训练多主体定制新框架。AnyMS利用文本提示、主体图像和布局约束三种输入,引入自底向上的双层注意力解耦机制,在生成过程中协调三者融合。全局解耦分离文本与视觉条件的交叉注意力,确保文本对齐;局部解耦将每个主体的注意力限制在其指定区域,防止主体冲突,从而保障身份保留与布局控制。此外,AnyMS采用预训练图像适配器提取与扩散模型对齐的主体特定特征,无需主体学习或适配器调优。大量实验表明,AnyMS达到当前最优性能,支持复杂组合,并可扩展至更多主体。

原文摘要 · Abstract (English)

Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorporate layout guidance to provide explicit spatial constraints. However, existing methods still struggle to balance three critical objectives: text alignment, subject identity preservation, and layout control, while the reliance on additional training further limits their scalability and efficiency. In this paper, we present AnyMS, a novel training-free framework for layout-guided multi-subject customization. AnyMS leverages three input conditions: text prompt, subject images, and layout constraints, and introduces a bottom-up dual-level attention decoupling mechanism to harmonize their integration during generation. Specifically, global decoupling separates cross-attention between textual and visual conditions to ensure text alignment. Local decoupling confines each subject's attention to its designated area, which prevents subject conflicts and thus guarantees identity preservation and layout control. Moreover, AnyMS employs pre-trained image adapters to extract subject-specific features aligned with the diffusion model, removing the need for subject learning or adapter tuning. Extensive experiments demonstrate that AnyMS achieves state-of-the-art performance, supporting complex compositions and scaling to a larger number of subjects.

多主体生成布局引导无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。