arXiv:2503.10365cs.CV2025-03被引 7

用局部视觉元素生成完整创意,让设计更贴近真实创作流程

Piece it Together: Part-Based Concepting with IP-Priors

  • 基于用户提供的碎片化视觉元素,自动补全缺失部分生成完整概念图
  • 在IP-Adapter+的特征空间上训练轻量级流匹配模型,实现上下文感知生成
  • 适配设计师、艺术创作者,尤其适合无文字描述的视觉构思场景

先进生成模型依赖文本条件生成图像,但视觉设计师常通过已有视觉片段(如独特翅膀、特定发型)激发灵感,这些片段仅构成完整概念的一部分。为此,我们提出一种生成框架,能将用户提供的部分视觉组件无缝整合为连贯构图,同时采样缺失部分以生成合理且完整的创意概念。该方法基于IP-Adapter+提取的强表征空间,训练了轻量级流匹配模型IP-Prior,利用领域特定先验实现多样化、上下文感知的生成。此外,我们提出一种基于LoRA的微调策略,显著提升IP-Adapter+在特定任务中的提示遵循能力,缓解其重建质量与提示契合度之间的固有权衡。

原文摘要 · Abstract (English)

Advanced generative models excel at synthesizing images but often rely on text-based conditioning. Visual designers, however, often work beyond language, directly drawing inspiration from existing visual elements. In many cases, these elements represent only fragments of a potential concept-such as an uniquely structured wing, or a specific hairstyle-serving as inspiration for the artist to explore how they can come together creatively into a coherent whole. Recognizing this need, we introduce a generative framework that seamlessly integrates a partial set of user-provided visual components into a coherent composition while simultaneously sampling the missing parts needed to generate a plausible and complete concept. Our approach builds on a strong and underexplored representation space, extracted from IP-Adapter+, on which we train IP-Prior, a lightweight flow-matching model that synthesizes coherent compositions based on domain-specific priors, enabling diverse and context-aware generations. Additionally, we present a LoRA-based fine-tuning strategy that significantly improves prompt adherence in IP-Adapter+ for a given task, addressing its common trade-off between reconstruction quality and prompt adherence.

图像生成视觉构思局部补全IP-Prior

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。