对比了遗传编程符号回归的初始化方法,发现初始优势很快消失。
Evaluation of Population Initialization Methods for Genetic Programming-based Symbolic Regression
- 用多种随机初始化和穷举符号回归优化解初始化
- 12个合成问题与1个真实数据集上无显著差异
- 适合关注初始化策略对进化算法影响的研究者
我们分析了在符号回归(SR)中优化遗传编程(GP)初始种群对解的准确性和复杂性的影响。比较了三种经典的随机初始化方法,以及使用穷举符号回归(ESR)获得的小规模优化解进行初始化。基于多目标进化算法NSGA-II的GP/SR实现,在十二个不同复杂度的合成问题和一个真实世界数据集上评估了每种初始化方法得到的最终帕累托前沿。结果表明,不同初始化方法在准确率或模型复杂度上均无显著差异。以ESR优化解初始化带来的初始优势在仅几代后即消失。研究显示,当初始种群多样性相似时,初始化方法对基于GP的符号回归最终帕累托前沿的影响可忽略不计。
原文摘要 · Abstract (English)
We analyze the effect of optimizing the initial population of genetic programming (GP) for symbolic regression (SR) on the accuracy and complexity of solutions. We compare three well-established random initialization methods as well as initialization with small optimized solutions from exhaustive symbolic regression (ESR) using a GP/SR implementation which is based on the multi-objective evolutionary algorithm NSGA-II. We compare the final Pareto fronts found with each initialization method on twelve synthetic problems of varying complexity and one real-world dataset. We find no significant differences in accuracy or model complexity among the initialization methods. The initial advantage of initialization with ESR disappears after only a few generations. Our results show that, given similar diversity in the initial population, the effect of the initialization method in GP-based symbolic regression on the final Pareto front is negligible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。