arXiv:2409.05305cs.LGcs.AI2024-09被引 2

用符号梯度找到神经网络隐空间中可读的数学公式

Closed-Form Interpretation of Neural Network Latent Spaces with Symbolic Gradients

  • 通过匹配神经元输入梯度,搜索人类可读的符号表达式
  • 成功从孪生网络中提取出矩阵不变量和动力系统守恒量
  • 适合对模型解释性、符号回归感兴趣的科研人员

研究表明,自编码器或孪生网络等人工神经网络在其隐空间中编码了有意义的概念。然而,缺乏无需先验知识即可将这些信息转化为人类可读形式的综合框架。在量化领域,概念通常以方程形式表达。为此,本文提出一种框架,用于寻找神经网络隐空间中神经元的闭式解释。该框架将训练好的神经网络嵌入到表示同一概念的函数等价类中,并通过符号搜索空间中的人类可读方程与该等价类求交,实现解释。计算上,该方法寻找一个符号表达式,其归一化梯度与特定神经元相对于输入变量的归一化梯度相匹配。实验表明,该方法能从孪生神经网络的隐空间中成功提取矩阵不变量和动力系统的守恒量。

原文摘要 · Abstract (English)

It has been demonstrated that artificial neural networks like autoencoders or Siamese networks encode meaningful concepts in their latent spaces. However, there does not exist a comprehensive framework for retrieving this information in a human-readable form without prior knowledge. In quantitative disciplines concepts are typically formulated as equations. Hence, in order to extract these concepts, we introduce a framework for finding closed-form interpretations of neurons in latent spaces of artificial neural networks. The interpretation framework is based on embedding trained neural networks into an equivalence class of functions that encode the same concept. We interpret these neural networks by finding an intersection between the equivalence class and human-readable equations defined by a symbolic search space. Computationally, this framework is based on finding a symbolic expression whose normalized gradients match the normalized gradients of a specific neuron with respect to the input variables. The effectiveness of our approach is demonstrated by retrieving invariants of matrices and conserved quantities of dynamical systems from latent spaces of Siamese neural networks.

神经网络解释符号回归隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。