arXiv:2505.23758cs.CV2025-05NeurIPS被引 13

无需训练即可用LoRA实现多概念图像编辑,像修图软件一样自由组合主体与风格。

LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers

  • 利用扩散模型中概念特征的早期空间一致性,生成可分离的潜在掩码。
  • 仅在目标概念区域融合对应LoRA权重,实现无缝多主题融合。
  • 无需重训练,适合创意设计、视觉叙事等快速迭代场景。

我们提出LoRAShop,首个基于LoRA模型的多概念图像编辑框架。该方法基于对Flux风格扩散变换器内部特征交互模式的观察:特定概念的变换器特征在去噪过程早期即激活空间上连贯的区域。利用这一特性,在前向传播中为每个概念生成解耦的潜在掩码,并仅在界定概念边界的区域内融合对应的LoRA权重。由此实现的编辑能将多个主体或风格自然融入原场景,同时保留全局上下文、光照和细节。实验表明,相比基线方法,LoRAShop在身份保持方面表现更优。通过消除重训练和外部约束,该方法使个性化扩散模型成为实用的「带LoRA的Photoshop」工具,为组合式视觉叙事和快速创意迭代开辟新路径。

原文摘要 · Abstract (English)

We introduce LoRAShop, the first framework for multi-concept image editing with LoRA models. LoRAShop builds on a key observation about the feature interaction patterns inside Flux-style diffusion transformers: concept-specific transformer features activate spatially coherent regions early in the denoising process. We harness this observation to derive a disentangled latent mask for each concept in a prior forward pass and blend the corresponding LoRA weights only within regions bounding the concepts to be personalized. The resulting edits seamlessly integrate multiple subjects or styles into the original scene while preserving global context, lighting, and fine details. Our experiments demonstrate that LoRAShop delivers better identity preservation compared to baselines. By eliminating retraining and external constraints, LoRAShop turns personalized diffusion models into a practical `photoshop-with-LoRAs' tool and opens new avenues for compositional visual storytelling and rapid creative iteration.

图像编辑LoRA扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。