arXiv:2505.20909cs.CV2025-05被引 4

让生成图像精准按布局、保留人物特征,还能自由创作任意内容。

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects

  • 用动态静态互补模块捕捉参考主体细节,不需微调模型。
  • 双布局控制机制提升训练与推理时的空间布局精度。
  • 适合需要高保真人物生成和灵活布局的创意设计场景。

扩散模型已显著推动文本到图像生成的发展,为个性化生成框架奠定了基础。然而,现有方法在布局控制精度上不足,且忽视了参考主体动态特征对生成质量的提升潜力。本文提出一种无需微调的布局可控个性化扩散模型(LCP-Diffusion),通过动态-静态互补视觉精炼模块全面捕捉参考主体的复杂细节,并引入双布局控制机制,在训练与推理阶段均实现稳健的空间控制。大量实验表明,该模型在身份保持与布局可控性方面表现优异。据我们所知,这是首个真正实现‘在任意位置生成任何内容’的开创性工作。

原文摘要 · Abstract (English)

Diffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our knowledge, this is a pioneering work enabling users to "create anything anywhere".

扩散模型个性化生成布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。