用最小描述长度原则实现符号回归的穷尽搜索,提升科学发现准确性。
(Exhaustive) Symbolic Regression and model selection by minimum description length
- 采用最小描述长度原理进行函数搜索与选择,直接量化复杂度与精度。
- 在三个天体物理难题中找到优于文献标准的函数表达式。
- 适用于科学建模,尤其适合需要可解释性模型的研究者。
符号回归是从未知数据中学习函数的机器学习方法。传统算法面临两大挑战:可能以高概率无法找到某个优秀函数;且在函数选择过程中存在歧义和未经验证的假设。为此,本文提出基于最小描述长度原则的穷尽搜索方法,使准确性和复杂度可在信息单位下直接权衡。该方法已公开实现,应用于三个天体物理开放问题:宇宙膨胀历史、星系中引力的有效行为以及暴胀场势能。在每项任务中,算法均发现多个优于现有文献标准的函数形式。这一通用方法在科学及其他领域具有广泛适用潜力。
原文摘要 · Abstract (English)
Symbolic regression is the machine learning method for learning functions from data. After a brief overview of the symbolic regression landscape, I will describe the two main challenges that traditional algorithms face: they have an unknown (and likely significant) probability of failing to find any given good function, and they suffer from ambiguity and poorly-justified assumptions in their function-selection procedure. To address these I propose an exhaustive search and model selection by the minimum description length principle, which allows accuracy and complexity to be directly traded off by measuring each in units of information. I showcase the resulting publicly available Exhaustive Symbolic Regression algorithm on three open problems in astrophysics: the expansion history of the universe, the effective behaviour of gravity in galaxies and the potential of the inflaton field. In each case the algorithm identifies many functions superior to the literature standards. This general purpose methodology should find widespread utility in science and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。