arXiv:2412.02126cs.LGcs.AI2024-12被引 5

对比8种常数优化方法,发现无通用最优方案

Benchmarking symbolic regression constant optimization schemes

  • 在进化搜索中测试8种常数优化方法
  • 不同方法在不同问题上表现差异显著
  • 提出树编辑距离评估符号准确性

符号回归是近年来发展迅速的机器学习技术,尤其在遗传编程(GPSR)方面进步明显。长期以来已知,在进化搜索中进行参数常数优化能显著提升性能,但不同研究采用的方法各异,尚无公认最优方案。本文在两个不同场景下,对十项经典基准问题评估了八种不同的参数优化方法。同时提出使用较少研究的树编辑距离(TED)作为符号准确性的度量指标,结合传统误差指标,建立综合模型性能分析框架。结果表明,不同常数优化方法在特定场景下表现更优,不存在适用于所有问题的全局最优方法。最后讨论了常见评估指标选择可能带来的偏差,导致某些方法看似表现更好。

原文摘要 · Abstract (English)

Symbolic regression is a machine learning technique, and it has seen many advancements in recent years, especially in genetic programming approaches (GPSR). Furthermore, it has been known for many years that constant optimization of parameters, during the evolutionary search, greatly increases GPSR performance However, different authors approach such tasks differently and no consensus exists regarding which methods perform best. In this work, we evaluate eight different parameter optimization methods, applied during evolutionary search, over ten known benchmark problems, in two different scenarios. We also propose using an under-explored metric called Tree Edit Distance (TED), aiming to identify symbolic accuracy. In conjunction with classical error measures, we develop a combined analysis of model performance in symbolic regression. We then show that different constant optimization methods perform better in certain scenarios and that there is no overall best choice for every problem. Finally, we discuss how common metric decisions may be biased and appear to generate better models in comparison.

符号回归常数优化遗传编程评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。