arXiv:2512.07355cs.AIcs.CV2025-12被引 2

用几何锥体统一监督与无监督概念学习,揭示两者本质相同。

A Geometric Unification of Concept Learning with Concept Cones

  • 将概念学习视为激活空间中的线性方向组合成锥体
  • 发现稀疏度与扩展因子存在最优平衡点,提升与人类概念对齐度
  • 为自编码器提供可量化的评估标准,适合模型可解释性研究者

解释性研究演化出两条路径:概念瓶颈模型(CBMs)通过人工标注定义概念,稀疏自编码器(SAEs)则通过稀疏编码自动发现概念。本文揭示二者均构建激活空间中的一组线性方向,其非负组合形成概念锥体。监督与无监督方法的区别在于如何选择该锥体。基于此,提出一个连接两者的框架:用CBM提供人类定义的几何基准,评估SAEs学习的锥体是否逼近或包含这些基准。该包含关系催生量化指标,关联了自编码器类型、稀疏度及扩展比等归纳偏置与合理概念(符合人类直觉但未必忠实于真实数据结构)的涌现。实验发现,在特定稀疏度和扩展因子下,几何与语义对齐达到最优。本工作通过共享几何框架统一两类方法,为评估SAE进展和概念合理性提供原则性工具。

原文摘要 · Abstract (English)

Two traditions of interpretability have evolved side by side but seldom spoken to each other: Concept Bottleneck Models (CBMs), which prescribe what a concept should be, and Sparse Autoencoders (SAEs), which discover what concepts emerge. While CBMs use supervision to align activations with human-labeled concepts, SAEs rely on sparse coding to uncover emergent ones. We show that both paradigms instantiate the same geometric structure: each learns a set of linear directions in activation space whose nonnegative combinations form a concept cone. Supervised and unsupervised methods thus differ not in kind but in how they select this cone. Building on this view, we propose an operational bridge between the two paradigms. CBMs provide human-defined reference geometries, while SAEs can be evaluated by how well their learned cones approximate or contain those of CBMs. This containment framework yields quantitative metrics linking inductive biases -- such as SAE type, sparsity, or expansion ratio -- to emergence of plausible\footnote{We adopt the terminology of \citet{jacovi2020towards}, who distinguish between faithful explanations (accurately reflecting model computations) and plausible explanations (aligning with human intuition and domain knowledge). CBM concepts are plausible by construction -- selected or annotated by humans -- though not necessarily faithful to the true latent factors that organise the data manifold.} concepts. Using these metrics, we uncover a ``sweet spot'' in both sparsity and expansion factor that maximizes both geometric and semantic alignment with CBM concepts. Overall, our work unifies supervised and unsupervised concept discovery through a shared geometric framework, providing principled metrics to measure SAE progress and assess how well discovered concept align with plausible human concepts.

可解释性概念学习几何建模自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。