从单图提取可组合的物体与属性概念,提升图像理解可解释性。
Intrinsic Concept Extraction Based on Compositional Interpretability
- 利用双曲空间建模概念层级与关系,实现精准解耦。
- 通过概念级优化保持复杂关系,确保概念可组合重建。
- 适合需要可解释性生成模型的研究者使用。
无监督概念提取旨在从单张图像中提取概念;然而现有方法难以提取可组合的内在概念。为此,本文提出新任务——可组合且可解释的内在概念提取(CI-ICE),旨在利用基于扩散的文生图模型,从单张图像中提取可组合的物体级与属性级概念,使原始概念可通过这些概念的组合重建。为达成目标,提出名为HyperExpress的方法,包含两个核心:首先,利用双曲空间固有的层次建模能力,实现概念精准解耦并保留概念间的层次结构与依赖关系;其次,引入概念级优化方法,映射概念嵌入空间以维持复杂的概念间关系,同时保障概念可组合性。实验表明,该方法在单图中提取可组合可解释的内在概念方面表现优异。
原文摘要 · Abstract (English)
Unsupervised Concept Extraction aims to extract concepts from a single image; however, existing methods suffer from the inability to extract composable intrinsic concepts. To address this, this paper introduces a new task called Compositional and Interpretable Intrinsic Concept Extraction (CI-ICE). The CI-ICE task aims to leverage diffusion-based text-to-image models to extract composable object-level and attribute-level concepts from a single image, such that the original concept can be reconstructed through the combination of these concepts. To achieve this goal, we propose a method called HyperExpress, which addresses the CI-ICE task through two core aspects. Specifically, first, we propose a concept learning approach that leverages the inherent hierarchical modeling capability of hyperbolic space to achieve accurate concept disentanglement while preserving the hierarchical structure and relational dependencies among concepts; second, we introduce a concept-wise optimization method that maps the concept embedding space to maintain complex inter-concept relationships while ensuring concept composability. Our method demonstrates outstanding performance in extracting compositionally interpretable intrinsic concepts from a single image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。