小模型能学会抽象代数结构,揭示神经网络隐含的数学规律。
Can Neural Networks Learn Small Algebraic Worlds? An Investigation Into the Group-theoretic Structures Learned By Narrow Models Trained To Predict Group Operations
- 用窄网络预测群运算,测试其是否捕捉代数结构。
- 模型能识别交换性与子群特征,但难学恒等元概念。
- 适合对神经网络可解释性与数学发现感兴趣的读者。
尽管现实数学研究常由问题驱动,但探索过程往往是开放式的,可能引出新的结构与洞见。这与机器学习中常见的应试式应用形成对比。为推进AI在数学中的作用,需超越简单问答。本文研究窄模型在固定任务(如预测群运算)中是否能学习更广泛的数学结构。通过一系列测试,评估模型是否掌握单位元、交换性或子群等概念。实验表明,模型能部分捕捉抽象代数属性:例如,在模运算中出现交换性迹象;且可训练线性分类器可靠区分某些子群元素(尽管训练数据未标注子群)。然而,未能有效提取单位元概念。结果表明,即使小型神经网络的表示也可能蕴含可挖掘的新数学结构。
原文摘要 · Abstract (English)
While a real-world research program in mathematics may be guided by a motivating question, the process of mathematical discovery is typically open-ended. Ideally, exploration needed to answer the original question will reveal new structures, patterns, and insights that are valuable in their own right. This contrasts with the exam-style paradigm in which the machine learning community typically applies AI to math. To maximize progress in mathematics using AI, we will need to go beyond simple question answering. With this in mind, we explore the extent to which narrow models trained to solve a fixed mathematical task learn broader mathematical structure that can be extracted by a researcher or other AI system. As a basic test case for this, we use the task of training a neural network to predict a group operation (for example, performing modular arithmetic or composition of permutations). We describe a suite of tests designed to assess whether the model captures significant group-theoretic notions such as the identity element, commutativity, or subgroups. Through extensive experimentation we find evidence that models learn representations capable of capturing abstract algebraic properties. For example, we find hints that models capture the commutativity of modular arithmetic. We are also able to train linear classifiers that reliably distinguish between elements of certain subgroups (even though no labels for these subgroups are included in the data). On the other hand, we are unable to extract notions such as the concept of the identity element. Together, our results suggest that in some cases the representations of even small neural networks can be used to distill interesting abstract structure from new mathematical objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。