揭示图像生成中关键语义的离散代码组合
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
- 通过样本级与代码本级分析,定位生成特定概念的离散令牌
- 在多个预训练VQGM模型上优于基线,解释效果清晰
- 适合需要可解释性或精准编辑的生成应用
向量量化生成模型(VQGMs)已成为强大的图像生成工具。然而,其核心组件——离散令牌的代码本——仍缺乏深入理解,例如哪些令牌对生成特定概念的图像至关重要?本文提出概念导向的令牌解释方法(CORTEX),通过识别概念相关的令牌组合来解释VQGMs。框架包含两种方法:(1)针对单张图像的样本级解释,分析其令牌重要性得分;(2)针对整个代码本的代码本级解释,挖掘全局相关令牌。实验表明,CORTEX在多个预训练VQGM模型上均优于基线,能清晰揭示生成过程中的令牌使用机制。该方法不仅提升VQGM的透明度,还可用于目标图像编辑与快捷特征检测。代码已开源。
原文摘要 · Abstract (English)
Vector-Quantized Generative Models (VQGMs) have emerged as powerful tools for image generation. However, the key component of VQGMs -- the codebook of discrete tokens -- is still not well understood, e.g., which tokens are critical to generate an image of a certain concept? This paper introduces Concept-Oriented Token Explanation (CORTEX), a novel approach for interpreting VQGMs by identifying concept-specific token combinations. Our framework employs two methods: (1) a sample-level explanation method that analyzes token importance scores in individual images, and (2) a codebook-level explanation method that explores the entire codebook to find globally relevant tokens. Experimental results demonstrate CORTEX's efficacy in providing clear explanations of token usage in the generative process, outperforming baselines across multiple pretrained VQGMs. Besides enhancing VQGMs transparency, CORTEX is useful in applications such as targeted image editing and shortcut feature detection. Our code is available at https://github.com/YangTianze009/CORTEX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。