发现激活函数对称性可导致表示离散化,影响模型可解释性。
Emergence of Quantised Representations Isolated to Anisotropic Functions
- 通过改变激活函数对称性,研究表示结构的演化机制。
- 离散对称性激活函数使表示离散化,连续对称性则保持连续。
- 该现象揭示了函数形式的隐含偏差,适合关注可解释性的研究者。
本文提出一种新方法,基于现有光点共振技术,通过仅改变激活函数的受控消融实验,探究自编码器中离散表示如何产生与组织。结果表明,当激活函数具有离散代数置换等变对称性时,表示趋向离散化;而连续正交等变定义下,表示保持连续。这验证了网络原语对称性可能携带未预期的归纳偏置,导致任务无关的表征结构。当代函数形式中的离散对称性被证明是预测离散表示形成的强指标,即量化效应。该发现提示应重新审视常见函数形式的潜在后果。此外,此机制支持一种离散表示生成的因果模型,可能是祖母细胞、离散编码、通用线性特征及某种超位置现象的前提。初步结果还显示,表示量化与重建误差显著上升相关,强化了其可能带来负面影响的猜想。
原文摘要 · Abstract (English)
Presented is a novel methodology for determining representational structure, which builds upon the existing Spotlight Resonance method. This new tool is used to gain insight into how discrete representations can emerge and organise in autoencoder models, through a controlled ablation study that alters only the activation function. Using this technique, the validity of whether function-driven symmetries can act as implicit inductive biases on representations is determined. Representations are found to tend to discretise when the activation functions are defined through a discrete algebraic permutation-equivariant symmetry. In contrast, they remain continuous under a continuous algebraic orthogonal-equivariant definition. This confirms the hypothesis that the symmetries of network primitives can carry unintended inductive biases, leading to task-independent artefactual structures in representations. The discrete symmetry of contemporary forms is shown to be a strong predictor for the production of symmetry-organised discrete representations emerging from otherwise continuous distributions -- a quantisation effect. This motivates further reassessment of functional forms in common usage due to such unintended consequences. Moreover, this supports a general causal model for a mode in which discrete representations may form, and could constitute a prerequisite for downstream interpretability phenomena, including grandmother neurons, discrete coding schemes, general linear features and a type of Superposition. Hence, this tool and proposed mechanism for the influence of functional form on representations may provide insights into interpretability research. Finally, preliminary results indicate that quantisation of representations correlates with a measurable increase in reconstruction error, reinforcing previous conjectures that this collapse can be detrimental.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。