arXiv:2501.01097cs.CV2025-01被引 37

让扩散模型精准控制图像中每个物体的位置和形状。

EliGen: Entity-Level Controlled Image Generation with Regional Attention

  • 用区域注意力机制实现无额外参数的实体级控制。
  • 在空间精度和图像质量上超越现有方法,支持任意形状掩码。
  • 可与多种开源模型结合,适合创意设计和精细编辑场景。

扩散模型在文本到图像生成方面取得显著进展,但仅依赖全局文本提示仍难以实现对图像中单个实体的细粒度控制。为此,我们提出EliGen——一种面向实体级可控图像生成的新框架。首先,提出无需额外参数的区域注意力机制,可无缝融合实体提示与任意形状的空间掩码;其次,构建了一个高质量数据集,包含细粒度的空间与语义实体标注,用于训练EliGen实现鲁棒准确的实体级操作,在空间精度和图像质量上均优于现有方法;此外,提出一种修复融合流水线,扩展其至多实体图像修复任务;最后,通过集成IP-Adapter、In-Context LoRA和MLLM等开源模型,验证其灵活性并开拓新创作可能。代码、模型及数据集已开源于https://github.com/modelscope/DiffSynth-Studio.git。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control over individual entities within an image. To address this limitation, we present EliGen, a novel framework for Entity-level controlled image Generation. Firstly, we put forward regional attention, a mechanism for diffusion transformers that requires no additional parameters, seamlessly integrating entity prompts and arbitrary-shaped spatial masks. By contributing a high-quality dataset with fine-grained spatial and semantic entity-level annotations, we train EliGen to achieve robust and accurate entity-level manipulation, surpassing existing methods in both spatial precision and image quality. Additionally, we propose an inpainting fusion pipeline, extending its capabilities to multi-entity image inpainting tasks. We further demonstrate its flexibility by integrating it with other open-source models such as IP-Adapter, In-Context LoRA and MLLM, unlocking new creative possibilities. The source code, model, and dataset are published at https://github.com/modelscope/DiffSynth-Studio.git.

图像生成扩散模型可控生成实体控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。