arXiv:2504.20197cs.LGcond-mat.dis-nn2025-04被引 2

用随机晶格模型解析神经网络特征,提升可解释性。

Representation Learning on a Random Lattice

  • 将数据分布视为随机晶格,从几何角度建模特征
  • 发现特征可分为上下文、组件和表面三类
  • 为可解释性研究提供新思路,适合关注模型透明度的读者

将深度神经网络学习到的表示分解为可解释特征,可显著提升其安全性和可靠性。为更好地理解特征,我们采用几何视角,将其视为映射嵌入数据分布的学得坐标系。我们提出将通用数据分布建模为随机晶格,并利用渗流理论分析其性质。学得特征被分类为上下文特征、组件特征和表面特征。该模型与近期机制可解释性研究结果定性一致,为未来研究指明方向。

原文摘要 · Abstract (English)

Decomposing a deep neural network's learned representations into interpretable features could greatly enhance its safety and reliability. To better understand features, we adopt a geometric perspective, viewing them as a learned coordinate system for mapping an embedded data distribution. We motivate a model of a generic data distribution as a random lattice and analyze its properties using percolation theory. Learned features are categorized into context, component, and surface features. The model is qualitatively consistent with recent findings in mechanistic interpretability and suggests directions for future research.

可解释性特征分解神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。