发现神经网络推理逻辑可被稀疏符号模式精确解释。
Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks
- 通过数学证明,两类通用准则导致符号模式涌现。
- 多数样本中符号交互具有高保真性与跨模型迁移能力。
- 为可解释学习提供新范式,适用于追求透明性的研究者。
人工神经网络(ANN)常被视为黑箱模型,可解释性是深度学习的核心挑战。尽管已有多种工程方法从特征归因和可视化等角度近似解释ANN,但其复杂推理逻辑能否被穷尽且简洁地表述为稀疏符号模式,仍是长期未解之问。本文揭示,在多种任务训练的广泛神经网络中,其推理逻辑确实可重构成稀疏符号交互。进一步理论证明,两类在各类任务中隐含存在的数学准则,导致了此类符号模式的涌现。实证表明,这两项准则在多数输入样本中成立。符号交互的保真性通过强样本间与模型间的可迁移性,以及对模型整体泛化能力的解释力得到验证。理论分析与大规模实验为神经网络的符号化解释奠定了基础,并揭示了符号模式对泛化能力的作用。研究还提出‘可沟通学习’新范式——直接在符号层面审视与调优推理逻辑,补充传统端到端学习。此外,符号模式的涌现提示:在特定条件下,其他黑箱系统也可能出现类似符号表示,因证明不依赖具体网络架构。
原文摘要 · Abstract (English)
Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning. Many engineering methods have been proposed to approximately explain the ANN from various perspectives, such as feature attribution and visualization. However, it remains a long-standing open question whether the complex inference logic of an ANN can be explained exhaustively and concisely as sparse symbolic patterns. This raises a deeper inquiry: does the emergence of symbolic patterns reflect a natural law rather than chance? Here, we show that across a broad class of ANNs trained on diverse tasks, their inference logic can indeed be reformulated as sparse symbolic interactions. We further prove that two common mathematical criteria, which are implicitly required across tasks, lead to the emergence of such sparse symbolic interactions. Empirical evidence confirms that the two criteria hold for the majority of input samples in diverse models. Furthermore, the faithfulness of these interactions is also demonstrated by their strong sample-to-sample and model-to-model transferability, as well as their ability to explain the overall generalization power of ANNs. Our theoretical analysis and extensive experiments provide a solid foundation for symbolic explanations of ANNs, and offer novel insights into the ANN's generalization power. Our findings also highlight the potential of communicative learning, a paradigm in which the inference logic of an ANN can be directly inspected and tuned at the level of symbolic patterns, thus complementing traditional end-to-end learning paradigm. Finally, the observed emergence of symbolic patterns in ANNs suggests that similar symbolic representations may also emerge in other types of black-box systems under certain conditions, because our proof does not depend on any specific ANN architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。