提出新方法PAN+SR,让符号回归可处理超多变量数据。
Ab Initio Nonparametric Variable Selection for Scalable Symbolic Regression with Large $p$
- 从零开始非参数筛选变量,降低搜索复杂度
- 在高维数据上提升19种符号回归方法性能
- 适合科学计算中变量超多且噪声大的场景
符号回归(SR)是一种强大的技术,用于发现描述数据中非线性关系的符号表达式,在可解释性、紧凑性和鲁棒性方面备受关注。然而,现有SR方法无法扩展到输入变量数量巨大的数据集(即极端规模符号回归),这在现代科学应用中十分常见。这种‘大p’情形常伴随测量误差,导致现有方法性能下降且表达式过于复杂难以解读。为解决这一可扩展性挑战,我们提出PAN+SR方法,将‘从头开始的非参数变量选择’与符号回归结合,高效预筛大规模输入空间,降低搜索复杂度的同时保持精度。非参数方法避免模型误设,支持一种称为参数辅助非参数(PAN)的策略。我们还扩展了开源基准平台SRBench,引入具有不同信噪比的高维回归问题。实验表明,PAN+SR持续提升了19种主流符号回归方法的性能,使其中若干方法在这些挑战性数据集上达到最新水平。
原文摘要 · Abstract (English)
Symbolic regression (SR) is a powerful technique for discovering symbolic expressions that characterize nonlinear relationships in data, gaining increasing attention for its interpretability, compactness, and robustness. However, existing SR methods do not scale to datasets with a large number of input variables (referred to as extreme-scale SR), which is common in modern scientific applications. This ``large $p$'' setting, often accompanied by measurement error, leads to slow performance of SR methods and overly complex expressions that are difficult to interpret. To address this scalability challenge, we propose a method called PAN+SR, which combines a key idea of ab initio nonparametric variable selection with SR to efficiently pre-screen large input spaces and reduce search complexity while maintaining accuracy. The use of nonparametric methods eliminates model misspecification, supporting a strategy called parametric-assisted nonparametric (PAN). We also extend SRBench, an open-source benchmarking platform, by incorporating high-dimensional regression problems with various signal-to-noise ratios. Our results demonstrate that PAN+SR consistently enhances the performance of 19 contemporary SR methods, enabling several to achieve state-of-the-art performance on these challenging datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。