让图像生成精准分离视觉概念,避免干扰
OmniPrism: Learning Disentangled Visual Concept for Image Generation
- 通过对比正交训练,从图像中解耦内容、风格等概念
- 构建20万对数据集,实现跨概念的精确分离
- 适合需要精准控制生成细节的创意设计场景
创意视觉概念生成常依赖参考图像中的特定概念以生成相关结果。然而现有方法通常仅支持单一维度的概念生成,或在多维度场景下易受无关概念干扰,导致概念混淆,阻碍创造性生成。为此,我们提出OmniPrism,一种面向创意图像生成的视觉概念解耦方法。该方法利用自然语言引导学习解耦的概念表征,并训练扩散模型融入这些概念。借助多模态提取器丰富的语义空间,实现从给定图像中对不同语义概念的解耦与指导。为解耦具有不同语义的概念,我们构建了包含20万对样本的配对概念解耦数据集(PCD-200K),每对样本共享同一概念(如内容、风格、构图)。通过对比正交解耦(COD)训练流程学习解耦表征,并将其注入扩散模型的额外交叉注意力层进行生成。设计块嵌入以适配扩散模型各模块的概念领域。大量实验表明,本方法能生成高质量、概念解耦且忠实于文本提示与目标概念的结果。
原文摘要 · Abstract (English)
Creative visual concept generation often draws inspiration from specific concepts in a reference image to produce relevant outcomes. However, existing methods are typically constrained to single-aspect concept generation or are easily disrupted by irrelevant concepts in multi-aspect concept scenarios, leading to concept confusion and hindering creative generation. To address this, we propose OmniPrism, a visual concept disentangling approach for creative image generation. Our method learns disentangled concept representations guided by natural language and trains a diffusion model to incorporate these concepts. We utilize the rich semantic space of a multimodal extractor to achieve concept disentanglement from given images and concept guidance. To disentangle concepts with different semantics, we construct a paired concept disentangled dataset (PCD-200K), where each pair shares the same concept such as content, style, and composition. We learn disentangled concept representations through our contrastive orthogonal disentangled (COD) training pipeline, which are then injected into additional diffusion cross-attention layers for generation. A set of block embeddings is designed to adapt each block's concept domain in the diffusion models. Extensive experiments demonstrate that our method can generate high-quality, concept-disentangled results with high fidelity to text prompts and desired concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。