arXiv:2503.21069cs.CV2025-03被引 1

高效生成多物体图像,解决布局不准与扩展难问题

Efficient Multi-Instance Generation with Janus-Pro-Dirven Prompt Parsing

  • 用轻量级模型解析文本生成布局,精准对齐语义与空间位置
  • 在COCO和LVIS上达到顶尖效果,仅需少量参数微调
  • 适合需要快速适配新场景的图像生成研究与应用

文本引导的扩散模型虽推动了条件图像生成,但在复杂多物体场景中仍面临空间定位不准和扩展性差的问题。本文提出两个核心模块:1)基于Janus-Pro的提示解析模块,通过仅10亿参数的紧凑架构,实现文本理解与布局生成的精准衔接;2)MIGLoRA,一种参数高效插件,将低秩适配(LoRA)集成至SD1.5与SD3的UNet和DiT主干网络中,保留基础模型参数,支持即插即用,实现高效微调。为全面评估,构建DescripBox与DescripBox-1024基准数据集,涵盖多样场景与分辨率。所提方法在COCO与LVIS基准上达当前最优性能,兼具布局保真度与开放世界合成的可扩展性。

原文摘要 · Abstract (English)

Recent advances in text-guided diffusion models have revolutionized conditional image generation, yet they struggle to synthesize complex scenes with multiple objects due to imprecise spatial grounding and limited scalability. We address these challenges through two key modules: 1) Janus-Pro-driven Prompt Parsing, a prompt-layout parsing module that bridges text understanding and layout generation via a compact 1B-parameter architecture, and 2) MIGLoRA, a parameter-efficient plug-in integrating Low-Rank Adaptation (LoRA) into UNet (SD1.5) and DiT (SD3) backbones. MIGLoRA is capable of preserving the base model's parameters and ensuring plug-and-play adaptability, minimizing architectural intrusion while enabling efficient fine-tuning. To support a comprehensive evaluation, we create DescripBox and DescripBox-1024, benchmarks that span diverse scenes and resolutions. The proposed method achieves state-of-the-art performance on COCO and LVIS benchmarks while maintaining parameter efficiency, demonstrating superior layout fidelity and scalability for open-world synthesis.

图像生成扩散模型多实例参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。