用扩散模型从单图中自动提取视觉概念,提升生成准确性
ICE: Intrinsic Concept Extraction from a Single Image via Diffusion Models
- 仅用文本到图像扩散模型,自动定位图像中的概念区域
- 将物体概念分解为内在概念与通用概念,实现细粒度解析
- 无需标注数据,适合研究视觉概念建模的学者使用
视觉概念定义的固有模糊性给基于扩散的文本到图像(T2I)模型从单张图像中准确学习概念带来挑战。现有方法缺乏系统性地可靠提取可解释内在概念的方法。为此,我们提出ICE(Intrinsic Concept Extraction),一种仅使用T2I模型从单张图像中自动、系统地提取内在概念的新框架。ICE包含两个关键阶段:第一阶段设计自动概念定位模块,识别图像中的文本相关概念及其对应掩码,为后续分析提供精准引导;第二阶段深入每个掩码,将物体级概念分解为内在概念与通用概念,实现更精细、可解释的视觉元素拆分。该框架在无监督条件下显著提升从单图中提取内在概念的性能。
原文摘要 · Abstract (English)
The inherent ambiguity in defining visual concepts poses significant challenges for modern generative models, such as the diffusion-based Text-to-Image (T2I) models, in accurately learning concepts from a single image. Existing methods lack a systematic way to reliably extract the interpretable underlying intrinsic concepts. To address this challenge, we present ICE, short for Intrinsic Concept Extraction, a novel framework that exclusively utilises a T2I model to automatically and systematically extract intrinsic concepts from a single image. ICE consists of two pivotal stages. In the first stage, ICE devises an automatic concept localization module to pinpoint relevant text-based concepts and their corresponding masks within the image. This critical stage streamlines concept initialization and provides precise guidance for subsequent analysis. The second stage delves deeper into each identified mask, decomposing the object-level concepts into intrinsic concepts and general concepts. This decomposition allows for a more granular and interpretable breakdown of visual elements. Our framework demonstrates superior performance on intrinsic concept extraction from a single image in an unsupervised manner. Project page: https://visual-ai.github.io/ice
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。