提出新方法提升符号回归的表达式分解能力与理论正确性
Deep Divide-and-Reduce in Symbolic Regression

- 基于数学推导构建表达式分解与化简的新框架
- 避免暴力搜索子结构,提升复杂方程求解成功率
- 适合需要高可靠性的科学建模与物理规律发现场景
符号回归(SR)旨在从数据中发现潜在规律并以数学表达式表示。现有机器学习方法往往缺乏对表达式内在数学与物理原理的深入理解。尽管开创性的AI Feynman方法利用数据中的数学特性,但其表达式简化机制适用范围窄,易在复杂方程上失败,且依赖于对子表达式的暴力搜索,严重限制了实用性。通过严格的数学推导与证明,我们提出深度分治化简符号回归方法(DDRSR)。该方法从根本上拓展了表达式分解与化简的适用性,避免了对子结构的暴力搜索,确保更广的通用性与严格的理论正确性。实证结果表明,这些理论优势在表达式分解与数值回归任务中均显著体现。最后,我们讨论了该范式的适用场景、内在局限及未来研究方向。
原文摘要 · Abstract (English)
Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and physical principles governing these expressions. While the pioneering AI Feynman method leverages the mathematical properties underlying the data, its expression simplification mechanism suffers from a narrow scope of applicability and is prone to failure on complex equations. Furthermore, its underlying mechanisms rely heavily on brute-force searches for sub-expressions, severely limiting its practical utility. Through rigorous mathematical deduction and proofs, we propose our method, Deep Divide and Reduce in Symbolic Regression (DDRSR). DDRSR fundamentally broadens the applicability of expression decomposition and reduction, circumvents the need for brute-force sub-structure searches, and ensures both wider versatility and strict theoretical correctness. Empirical evaluations demonstrate that these theoretical principles yield significant advantages in both expression decomposition and numerical regression tasks. Finally, we discuss the applicable scenarios and inherent limitations of this paradigm, alongside promising directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。