用扩散模型从一张图中分解出定制化视觉概念,生成高质量图像。
CusConcept: Customized Visual Concept Decomposition with Diffusion Models
- 分两阶段:先构建人类指定轴上的概念词典,再优化生成质量。
- 能从单图生成多视角的高质量视觉概念图像,支持文本生成提示。
- 适合需要可控图像生成与概念拆解的研究者或设计师使用。
让生成模型从单张图像中分解出视觉概念是一项复杂且具有挑战性的任务。本文提出新任务——定制化概念分解,目标是利用扩散模型从单张图像中分解出不同视角的视觉概念。为此,我们设计了两阶段框架CusConcept(定制化视觉概念分解),提取可嵌入文本生成提示的定制化视觉概念嵌入向量。第一阶段采用词汇引导的概念分解机制,在人类指定的概念轴上构建词汇库,通过检索对应词汇并学习锚点权重获得分解概念。第二阶段进行联合概念优化,提升生成图像的保真度和质量。我们还构建了一个评估基准,用于衡量开放世界概念分解任务的表现。实验表明,该方法能有效生成高质量的分解概念图像,并产出相关词汇预测作为辅助结果。大量定性与定量实验验证了CusConcept的有效性。
原文摘要 · Abstract (English)
Enabling generative models to decompose visual concepts from a single image is a complex and challenging problem. In this paper, we study a new and challenging task, customized concept decomposition, wherein the objective is to leverage diffusion models to decompose a single image and generate visual concepts from various perspectives. To address this challenge, we propose a two-stage framework, CusConcept (short for Customized Visual Concept Decomposition), to extract customized visual concept embedding vectors that can be embedded into prompts for text-to-image generation. In the first stage, CusConcept employs a vocabulary-guided concept decomposition mechanism to build vocabularies along human-specified conceptual axes. The decomposed concepts are obtained by retrieving corresponding vocabularies and learning anchor weights. In the second stage, joint concept refinement is performed to enhance the fidelity and quality of generated images. We further curate an evaluation benchmark for assessing the performance of the open-world concept decomposition task. Our approach can effectively generate high-quality images of the decomposed concepts and produce related lexical predictions as secondary outcomes. Extensive qualitative and quantitative experiments demonstrate the effectiveness of CusConcept.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。