arXiv:2507.10678cs.LGcs.AI2025-07

用群论分析进位机制,揭示神经网络学对称性的关键因素

A Group Theoretic Analysis of the Symmetries Underlying Base Addition and Their Learnability by Neural Networks

  • 用群论解析进位函数的数学结构,发现多种可替代进位方式
  • 简单网络在合适输入和进位规则下实现远超训练范围的泛化
  • 结果揭示神经网络学习对称性的内在偏好,对认知与模型设计有启发

神经网络在建模人类认知和人工智能中面临的核心挑战之一,是设计能高效学习支持极端泛化的函数系统。其根本在于发现并实现对称性函数。本文以进位加法这一典型对称性任务为研究对象,通过群论分析进位机制——即当和超过基数时,将余数传递到更高位的规则。分析揭示了同一基数下存在多种可选的进位函数,并引入量化指标进行表征。随后,我们通过训练神经网络在不同进位规则下执行进位加法,比较其学习效率与性能,探究网络结构对对称性学习的归纳偏置影响。结果表明,即使简单网络,在合适的输入格式和进位函数下也能实现极端泛化,且学习能力与进位函数的结构密切相关。该研究对认知科学与机器学习具有重要意义。

原文摘要 · Abstract (English)

A major challenge in the use of neural networks both for modeling human cognitive function and for artificial intelligence is the design of systems with the capacity to efficiently learn functions that support radical generalization. At the roots of this is the capacity to discover and implement symmetry functions. In this paper, we investigate a paradigmatic example of radical generalization through the use of symmetry: base addition. We present a group theoretic analysis of base addition, a fundamental and defining characteristic of which is the carry function -- the transfer of the remainder, when a sum exceeds the base modulus, to the next significant place. Our analysis exposes a range of alternative carry functions for a given base, and we introduce quantitative measures to characterize these. We then exploit differences in carry functions to probe the inductive biases of neural networks in symmetry learning, by training neural networks to carry out base addition using different carries, and comparing efficacy and rate of learning as a function of their structure. We find that even simple neural networks can achieve radical generalization with the right input format and carry function, and that learnability is closely correlated with carry function structure. We then discuss the relevance this has for cognitive science and machine learning.

群论对称性神经网络泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。