arXiv:2607.10420cs.SEcs.LG2026-07

模型不稳定不是噪声,而是可测量管理的特性。

Is Model Instability just Noise to be Tolerated or a Property that can be Managed?

论文配图:Is Model Instability just Noise to be Tolerated or a Property that can be Managed?
图 1 · 摘自论文原文
  • 通过调整标签使用、模型复杂度和评分策略提升稳定性
  • 模型一致率提高4.8倍,误差标准差下降22%
  • 适合关注模型可信度与优化稳定性的软件工程研究者

在软件工程优化中,同一分析重复运行常得不同模型与结论,削弱信任。我们在127个多目标优化问题(共12,700个测试用例)中发现,即使使用先进优化器,重复运行仅在13.7%的案例上达成一致。我们主张,这种不稳定性不应被视为需容忍的噪声,而应被量化与管理。通过调整标签分配、模型复杂度及分割评分方式,新配置使模型一致性提升4.8倍。平均优化误差标准差从17.4降至13.6(降幅22%),推荐质量不降反升。在127个数据集上,新设置在质量上统计显著优于默认配置的有119个,而默认仅74个。进一步测试因果干预与数据局部性方法,效果有限,表明存在稳定性下限。我们认为,数据本身的噪声、标签稀缺、代理目标及大量近似等效模型是根本限制。因此,建议将不稳定性作为标准评估维度,与性能一同报告,并用于校准对单次运行结果的信任。本文方法为未来减少SBSE不稳定性提供基准。为支持开放科学,提供复现包:https://tinyurl.com/Model-Instability

原文摘要 · Abstract (English)

In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability

软件工程模型稳定优化可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。