算法选择评估方法存在漏洞,可能误导研究方向
The Pitfalls of Benchmarking in Algorithm Selection: What We Are Getting Wrong
- 指出'留实例外'评估法会误判无意义特征和模型的性能
- 发现目标函数尺度敏感指标会导致元模型评估结果虚高
- 提醒研究者需严谨设计评估流程,避免无效投入
算法选择旨在为给定问题找到最优算法,在连续黑箱优化中具有关键作用。常见方法是用一组特征表示优化函数,再训练机器学习元模型进行算法选择。尽管多种方法已证明其有效性,但并非所有评估方式都适用于元模型性能评估。本文揭示了领域内常见的方法学问题:首先,'留实例外'评估技术存在缺陷,非信息性特征与元模型也能获得高准确率,这在合理评估框架下不应出现;其次,当使用对目标函数尺度敏感的指标衡量优化算法性能时,需谨慎考虑其对元模型构建、预测结果及误差分析的影响,此类指标可能造成元模型性能被过度乐观地评估。本文强调评估严谨性的重要性,松散的方法论可能导致研究误导、资源浪费并引入噪声。
原文摘要 · Abstract (English)
Algorithm selection, aiming to identify the best algorithm for a given problem, plays a pivotal role in continuous black-box optimization. A common approach involves representing optimization functions using a set of features, which are then used to train a machine learning meta-model for selecting suitable algorithms. Various approaches have demonstrated the effectiveness of these algorithm selection meta-models. However, not all evaluation approaches are equally valid for assessing the performance of meta-models. We highlight methodological issues that frequently occur in the community and should be addressed when evaluating algorithm selection approaches. First, we identify flaws with the "leave-instance-out" evaluation technique. We show that non-informative features and meta-models can achieve high accuracy, which should not be the case with a well-designed evaluation framework. Second, we demonstrate that measuring the performance of optimization algorithms with metrics sensitive to the scale of the objective function requires careful consideration of how this impacts the construction of the meta-model, its predictions, and the model's error. Such metrics can falsely present overly optimistic performance assessments of the meta-models. This paper emphasizes the importance of careful evaluation, as loosely defined methodologies can mislead researchers, divert efforts, and introduce noise into the field
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。