arXiv:2604.24936cs.LGstat.ML2026-04被引 1

提出统一理论框架,让无监督概念提取更可靠。

A Unifying Framework for Unsupervised Concept Extraction

  • 将概念提取视为识别生成模型,统一分析视角。
  • 给出通用可辨识性定理,简化证明复杂度。
  • 适合研究可解释性与模型控制的学者参考。

概念提取技术(如稀疏自编码器和转换器)旨在从低层非符号表示中提取高层符号概念。当这些概念用于下游任务(如模型调控和遗忘)时,理解其保证与否至关重要。本文提出一个统一的理论框架,将概念提取任务建模为识别生成模型。我们提出一个通用的可辨识性元定理,将建立可辨识性保证的问题简化为两个集合交集的刻画问题。通过在多种广泛应用的方法上验证,该元定理显著简化了证明过程,为发展新的、基于原理的概念提取方法铺平道路。

原文摘要 · Abstract (English)

Techniques for concept extraction, such as sparse autoencoders and transcoders, aim to extract high-level symbolic concepts from low-level nonsymbolic representations. When these extracted concepts are used for downstream tasks such as model steering and unlearning, it is essential to understand their guarantees, or lack thereof. In this work, we present a unified theoretical framework for unsupervised concept extraction, in which we frame the task of concept extraction as identifying a generative model. We present a general meta-theorem for identifiability, which reduces the problem of establishing identifiability guarantees to the problem of characterizing the intersection of two sets. As we demonstrate on a range of widely-used approaches, this meta-theorem substantially simplifies the task of proving such guarantees, thus paving the way for the development of new, principled approaches for concept extraction.

概念提取理论分析可辨识性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。