arXiv:2505.03977cs.LGcs.NE2025-05中稿 · Genetic and Evolut…被引 16

更新符号回归基准,推动算法公平对比与可持续演进

Call for Action: towards the next generation of symbolic regression benchmark

  • 扩展SRBench基准,方法数近翻倍,优化评估指标与可视化
  • 发现无单一算法在所有数据集上占优,复杂度、精度与能耗存在权衡
  • 呼吁社区共建动态基准,倡导标准化调参与节能实现

符号回归(SR)是一种强大的技术,用于发现可解释的数学表达式。然而,由于算法、数据集和评估标准的多样性,对SR方法进行基准测试仍具挑战性。本文呈现了SRBench的更新版本:方法数量几乎翻倍,评估指标得到优化,结果可视化更完善。我们还分析了模型复杂度、准确率与能耗之间的权衡。结果表明,没有单一算法在所有数据集上表现最优。我们呼吁符号回归社区共同维护并持续演进SRBench,使其成为反映该领域最新进展的动态基准,通过标准化超参数调优、执行约束和计算资源分配来提升可比性。同时提出淘汰标准以保持基准相关性,并讨论改进算法的最佳实践,如自适应超参数调优和能效优化实现。

原文摘要 · Abstract (English)

Symbolic Regression (SR) is a powerful technique for discovering interpretable mathematical expressions. However, benchmarking SR methods remains challenging due to the diversity of algorithms, datasets, and evaluation criteria. In this work, we present an updated version of SRBench. Our benchmark expands the previous one by nearly doubling the number of evaluated methods, refining evaluation metrics, and using improved visualizations of the results to understand the performances. Additionally, we analyze trade-offs between model complexity, accuracy, and energy consumption. Our results show that no single algorithm dominates across all datasets. We propose a call for action from SR community in maintaining and evolving SRBench as a living benchmark that reflects the state-of-the-art in symbolic regression, by standardizing hyperparameter tuning, execution constraints, and computational resource allocation. We also propose deprecation criteria to maintain the benchmark's relevance and discuss best practices for improving SR algorithms, such as adaptive hyperparameter tuning and energy-efficient implementations.

符号回归基准测试算法评估能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。