高精度模型未必带来更好调优结果,研究揭示精度与调优效果并非线性相关。
Accuracy Can Lie: On the Impact of Surrogate Model in Configuration Tuning
- 通过大规模实证研究验证不同模型精度对调优效果的影响。
- 超58%情况下更高精度反而无提升,24%导致性能下降。
- 提醒社区警惕‘精度即一切’的误区,关注模型适用性而非单纯精度。
为降低配置调优中昂贵的测量成本,通常会构建代理模型替代真实系统进行性能评估。然而,普遍认为模型精度越高,调优效果越好,这一‘精度即一切’的观念促使研究者不断追求更高精度模型,并批评调优器因模型不准确而失败。但该做法引发新问题:现有工作报告的小幅精度提升是否真对调优有显著影响?模型精度在调优质量中起何作用?为此,我们开展了迄今最大规模的实证研究,持续13个月、全天候运行,涵盖10种模型、17种调优器和29个系统,在四种常用指标下共分析13,612个案例。惊人发现:精度可能欺骗人——在某些设置下,高达58%的情况中更高精度未带来调优改善,甚至24%导致性能下降。还发现多数调优器所选模型非最优,且显著提升调优质量所需精度变化量随模型精度区间而异。基于适应度景观分析,深入探讨了背后机制,提出若干经验教训与未来方向。最重要的是,本研究向社区发出明确信号:应从‘精度即一切’的惯性思维中退一步。
原文摘要 · Abstract (English)
To ease the expensive measurements during configuration tuning, it is natural to build a surrogate model as the replacement of the system, and thereby the configuration performance can be cheaply evaluated. Yet, a stereotype therein is that the higher the model accuracy, the better the tuning result would be. This "accuracy is all" belief drives our research community to build more and more accurate models and criticize a tuner for the inaccuracy of the model used. However, this practice raises some previously unaddressed questions, e.g., Do those somewhat small accuracy improvements reported in existing work really matter much to the tuners? What role does model accuracy play in the impact of tuning quality? To answer those related questions, we conduct one of the largest-scale empirical studies to date-running over the period of 13 months 24*7-that covers 10 models, 17 tuners, and 29 systems from the existing works while under four different commonly used metrics, leading to 13,612 cases of investigation. Surprisingly, our key findings reveal that the accuracy can lie: there are a considerable number of cases where higher accuracy actually leads to no improvement in the tuning outcomes (up to 58% cases under certain setting), or even worse, it can degrade the tuning quality (up to 24% cases under certain setting). We also discover that the chosen models in most proposed tuners are sub-optimal and that the required % of accuracy change to significantly improve tuning quality varies according to the range of model accuracy. Deriving from the fitness landscape analysis, we provide in-depth discussions of the rationale behind, offering several lessons learned as well as insights for future opportunities. Most importantly, this work poses a clear message to the community: we should take one step back from the natural "accuracy is all" belief for model-based configuration tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。