用五次方程根分类测试机器学习能否自动发现可解释数学规律
On the Limits of Interpretable Machine Learning in Quintic Root Classification
- 给定关键点符号变化特征后,决策树可达到与神经网络相当的准确率
- 神经网络在分布内达84.3%准确率,但依赖连续几何逼近而非符号规则
- 97.5%决策结构由单一不变量决定,说明可解释性需显式结构先验
机器学习能否从原始数值数据中自主恢复可解释的数学结构?我们以五次多项式实根配置分类为结构化基准进行研究。测试了决策树、逻辑回归、支持向量机、随机森林、梯度提升、XGBoost、符号回归和神经网络等模型。神经网络仅用原始系数即达84.3%±0.9%平衡准确率,而决策树表现较差(59.9%±0.9%)。但当引入关键点符号变化的显式特征后,决策树性能提升至84.2%±1.2%,并生成明确分类规则。知识蒸馏显示该单一不变量解释了97.5%的决策结构。分布外、数据效率和噪声鲁棒性分析表明,神经网络学习的是连续、数据依赖的几何近似,而非尺度不变的符号规则。这一几何逼近与符号不变性的区别,解释了各模型预测性能与可解释性间的差距。尽管高精度可实现,但未发现所测模型能从原始系数中自主恢复离散的人类可读数学规则。结果表明,在结构化数学领域,可解释性可能需要显式结构归纳偏置,而非纯数据驱动逼近。
原文摘要 · Abstract (English)
Can Machine Learning (ML) autonomously recover interpretable mathematical structure from raw numerical data? We aim to answer this question using the classification of real-root configurations of polynomials up to degree five as a structured benchmark. We tested an extensive set of ML models, including decision trees, logistic regression, support vector machines, random forest, gradient boosting, XGBoost, symbolic regression, and neural networks. Neural networks achieved strong in-distribution performance on quintic classification using raw coefficients alone (84.3% + or - 0.9% balanced accuracy), whereas decision trees perform substantially worse (59.9% + or - 0.9\%). However, when provided with an explicit feature capturing sign changes at critical points, decision trees match neural performance (84.2% + or - 1.2%) and yield explicit classification rules. Knowledge distillation reveals that this single invariant accounts for 97.5% of the extracted decision structure. Out-of-distribution, data-efficiency, and noise robustness analyses indicate that neural networks learn continuous, data-dependent geometric approximations of the decision boundary rather than recovering scale-invariant symbolic rules. This distinction between geometric approximation and symbolic invariance explains the gap between predictive performance and interpretability observed across models. Although high predictive accuracy is attainable, we find no evidence that the evaluated ML models autonomously recover discrete, human-interpretable mathematical rules from raw coefficients. These results suggest that, in structured mathematical domains, interpretability may require explicit structural inductive bias rather than purely data-driven approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。