arXiv:2508.14440cs.CV2025-08ICCV被引 4

让多个物体精准按布局生成,保持身份一致

MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion

  • 用显式语义扩展融合布局与文本引导
  • 零样本端到端生成,空间精度和身份一致性更优
  • 适合需要精细多对象合成的视觉设计场景

现有文本到图像扩散模型在高质量图像生成方面表现卓越,但实现具有精确空间控制的多主体组合合成仍是重大挑战。本文针对布局可控的多主体合成(LMS)任务,要求既忠实还原参考主体,又将其准确放置于指定区域。尽管近期进展分别提升了布局控制与主体合成能力,现有方法难以同时满足空间精度与身份保留的双重需求。为此,我们提出MUSE统一合成框架,采用拼接交叉注意力(CCA)机制,通过显式语义空间扩展将布局信息与文本引导无缝融合。该机制实现了空间约束与文本描述间的双向模态对齐,且无干扰。此外,设计渐进式两阶段训练策略,将LMS任务分解为可学习的子目标以实现有效优化。大量实验表明,MUSE在零样本端到端生成中,相比现有方案具备更优的空间准确性和身份一致性,推动可控图像合成的前沿发展。代码与模型已公开于https://github.com/pf0607/MUSE。

原文摘要 · Abstract (English)

Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a significant challenge. In this work, we address the task of layout-controllable multi-subject synthesis (LMS), which requires both faithful reconstruction of reference subjects and their accurate placement in specified regions within a unified image. While recent advancements have separately improved layout control and subject synthesis, existing approaches struggle to simultaneously satisfy the dual requirements of spatial precision and identity preservation in this composite task. To bridge this gap, we propose MUSE, a unified synthesis framework that employs concatenated cross-attention (CCA) to seamlessly integrate layout specifications with textual guidance through explicit semantic space expansion. The proposed CCA mechanism enables bidirectional modality alignment between spatial constraints and textual descriptions without interference. Furthermore, we design a progressive two-stage training strategy that decomposes the LMS task into learnable sub-objectives for effective optimization. Extensive experiments demonstrate that MUSE achieves zero-shot end-to-end generation with superior spatial accuracy and identity consistency compared to existing solutions, advancing the frontier of controllable image synthesis. Our code and model are available at https://github.com/pf0607/MUSE.

图像生成多主体合成布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。