arXiv:2605.05223cs.LGcs.AI2026-05

揭示特征组合在稀疏自编码器中的结构不稳定性

Structural Instability of Feature Composition

  • 用高维稀疏锥流形建模激活空间,分析组合激活的几何不稳定性
  • 发现特征组合在高偏差下会因ReLU导致系统性干扰累积
  • 适用于研究大模型可解释性与特征控制的学者

稀疏自编码器(SAEs)已成为解耦基于Transformer架构中特征超叠加现象的强大工具,支持通过激活操控实现精确控制。然而,关于组合操控——即同时激活不同语义潜变量——的理论基础仍不明确。主流线性表示假设忽略了过完备字典中出现的非线性干扰效应。本文提出一种几何框架,用于分析特征联合的不稳定性。将激活空间建模为高维稀疏锥流形,推导出在球形字典模型下的渐近组合坍缩阈值,其由信号锥的高斯均宽(统计维度)决定。进一步表明,在高偏置状态下,ReLU整流将微小相关性引起的方差波动转化为系统性漂移,随组合不断累积,产生符合棘轮效应的干扰增长。在CLEVR中提取的结构化语义特征上验证了预测的尺度趋势,其中层级相关性加速了过渡,相比随机基线更明显。结果揭示了基于联合操控的可扩展性几何限制,并呼吁设计超越简单线性叠加的组合机制以主动管理干扰。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have emerged as a powerful paradigm for disentangling feature superposition in transformer-based architectures, enabling precise control via activation steering. However, the theoretical foundations of compositional steering -- the simultaneous activation of distinct semantic latents -- remain under-explored. The prevailing Linear Representation Hypothesis often abstracts away non-linear interference effects that arise in overcomplete dictionaries. We present a geometric framework for analyzing the instability of feature unions. Modeling the activation space as a high-dimensional sparse cone manifold, we derive an asymptotic compositional-collapse threshold under a spherical dictionary model, characterized by the Gaussian mean width (statistical dimension) of the signal cone. We further show that, in the high-bias regime, ReLU rectification converts microscopic correlation-induced variance fluctuations into a systematic drift that accumulates under composition, yielding interference growth consistent with a ratchet effect. We validate the predicted scaling trends on structured semantic features extracted from CLEVR, where hierarchical correlations accelerate the transition relative to random baselines. Together, our results highlight geometric constraints on the scalability of union-based steering and motivate composition mechanisms that explicitly manage interference beyond naive linear superposition.

稀疏编码特征组合可解释性神经网络几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。