arXiv:2506.19708cs.GRcs.AI2025-06被引 5

用稀疏自编码器发现生成模型忽略的常见概念

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

  • 用稀疏自编码器提取可解释的概念嵌入,量化真实与生成图像中概念的差异
  • 在32,000个概念上训练最大规模SAE,发现模型对鸟食器、光盘等概念抑制明显
  • 识别出木纹、棕榈树等被过度生成的夸张概念,以及训练数据复现的模板化错误

尽管大规模训练的生成图像模型表现优异,但常无法生成看似简单却应存在于训练数据中的概念,如人手或四件物品组合。这些失败模式多为零散记录,难以判断是偶然异常还是模型结构性缺陷。为此,我们提出系统方法,识别并刻画“概念盲区”——即训练数据中存在但在生成结果中缺失或失真的概念。方法基于稀疏自编码器(SAEs)提取可解释的概念嵌入,实现真实与生成图像间概念出现频率的定量对比。我们在DINOv2特征上训练了包含32,000个概念的原型SAE(RA-SAE),为迄今最大规模此类模型,支持细粒度分析概念偏差。应用于Stable Diffusion 1.5/2.1、PixArt和Kandinsky四个主流生成模型,揭示特定受抑制盲区(如鸟食器、DVD光盘、文档空白处)及被夸大的盲区(如木纹背景、棕榈树)。在单样本层面,进一步分离出记忆化伪影——模型复现训练中见过的特定视觉模板。整体提出一种理论扎实的框架,通过评估生成模型与数据生成过程的概念保真度,系统识别其概念盲区。

原文摘要 · Abstract (English)

Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands or objects appearing in groups of four -- that are reasonably expected to appear in the training data. These failure modes have largely been documented anecdotally, leaving open the question of whether they reflect idiosyncratic anomalies or more structural limitations of these models. To address this, we introduce a systematic approach for identifying and characterizing "conceptual blindspots" -- concepts present in the training data but absent or misrepresented in a model's generations. Our method leverages sparse autoencoders (SAEs) to extract interpretable concept embeddings, enabling a quantitative comparison of concept prevalence between real and generated images. We train an archetypal SAE (RA-SAE) on DINOv2 features with 32,000 concepts -- the largest such SAE to date -- enabling fine-grained analysis of conceptual disparities. Applied to four popular generative models (Stable Diffusion 1.5/2.1, PixArt, and Kandinsky), our approach reveals specific suppressed blindspots (e.g., bird feeders, DVD discs, and whitespaces on documents) and exaggerated blindspots (e.g., wood background texture and palm trees). At the individual datapoint level, we further isolate memorization artifacts -- instances where models reproduce highly specific visual templates seen during training. Overall, we propose a theoretically grounded framework for systematically identifying conceptual blindspots in generative models by assessing their conceptual fidelity with respect to the underlying data-generating process.

生成模型概念盲区稀疏自编码器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。