超参数调优会加剧模型预测不一致,影响高风险决策可靠性。
The Role of Hyperparameters in Predictive Multiplicity
- 分析六种表格式数据模型的超参数对预测分歧的影响。
- 调优提升性能但显著增加预测差异,极端梯度提升差异最大。
- 揭示性能与稳定性的权衡,警示决策中潜在的任意性风险。
本文研究超参数在预测多重性中的关键作用,即相同数据训练的不同模型对相同输入产生不同预测的现象,可能严重影响信用评估、招聘和医疗诊断等高风险决策。针对六种广泛使用的表格数据模型——弹性网络、决策树、k近邻、支持向量机、随机森林和极端梯度提升,我们探讨超参数调优如何影响预测分歧程度,基于21个基准数据集的实验表明,如弹性网络中的lambda、支持向量机中的gamma、极端梯度提升中的alpha等关键超参数显著影响预测一致性。调优虽带来性能提升,却也导致预测差异增大,其中极端梯度提升表现出最高分歧和严重不稳定。这揭示了性能优化与预测一致性之间的权衡,提示存在预测结果随意化的风险。尽管预测多重性有助于实现公平性等特定目标并减少对单一模型的依赖,但也使决策过程复杂化,可能导致任意或无依据的结果。
原文摘要 · Abstract (English)
This paper investigates the critical role of hyperparameters in predictive multiplicity, where different machine learning models trained on the same dataset yield divergent predictions for identical inputs. These inconsistencies can seriously impact high-stakes decisions such as credit assessments, hiring, and medical diagnoses. Focusing on six widely used models for tabular data - Elastic Net, Decision Tree, k-Nearest Neighbor, Support Vector Machine, Random Forests, and Extreme Gradient Boosting - we explore how hyperparameter tuning influences predictive multiplicity, as expressed by the distribution of prediction discrepancies across benchmark datasets. Key hyperparameters such as lambda in Elastic Net, gamma in Support Vector Machines, and alpha in Extreme Gradient Boosting play a crucial role in shaping predictive multiplicity, often compromising the stability of predictions within specific algorithms. Our experiments on 21 benchmark datasets reveal that tuning these hyperparameters leads to notable performance improvements but also increases prediction discrepancies, with Extreme Gradient Boosting exhibiting the highest discrepancy and substantial prediction instability. This highlights the trade-off between performance optimization and prediction consistency, raising concerns about the risk of arbitrary predictions. These findings provide insight into how hyperparameter optimization leads to predictive multiplicity. While predictive multiplicity allows prioritizing domain-specific objectives such as fairness and reduces reliance on a single model, it also complicates decision-making, potentially leading to arbitrary or unjustified outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。