用贝叶斯方法提升符号回归抗噪能力,发现更准确的物理方程。
Bayesian Symbolic Regression via Posterior Sampling
- 基于序列蒙特卡洛框架,通过后验采样搜索符号表达式
- 在噪声数据上表现更优,过拟合减少,泛化能力更强
- 适合需要可解释方程的科学发现与工程设计场景
符号回归能直接从数据中发现控制方程,但对噪声敏感限制了其应用。本文提出一种基于序列蒙特卡洛(SMC)的贝叶斯符号回归框架,通过近似符号表达式的后验分布,提升抗噪性并实现不确定性量化。不同于传统遗传编程,该算法结合概率选择、自适应退火和归一化边缘似然,高效探索符号表达式空间,获得更简洁且泛化能力更强的表达式。相比标准遗传编程基线,在具有挑战性的噪声基准数据集上表现更优,显著降低过拟合倾向,提升方程发现的准确性与可解释性,为科学发现与工程设计中的鲁棒符号回归提供新路径。
原文摘要 · Abstract (English)
Symbolic regression is a powerful tool for discovering governing equations directly from data, but its sensitivity to noise hinders its broader application. This paper introduces a Sequential Monte Carlo (SMC) framework for Bayesian symbolic regression that approximates the posterior distribution over symbolic expressions, enhancing robustness and enabling uncertainty quantification for symbolic regression in the presence of noise. Differing from traditional genetic programming approaches, the SMC-based algorithm combines probabilistic selection, adaptive tempering, and the use of normalized marginal likelihood to efficiently explore the search space of symbolic expressions, yielding parsimonious expressions with improved generalization. When compared to standard genetic programming baselines, the proposed method better deals with challenging, noisy benchmark datasets. The reduced tendency to overfit and enhanced ability to discover accurate and interpretable equations paves the way for more robust symbolic regression in scientific discovery and engineering design applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。