通过模拟分类模型预测,评估其对优化结果的影响。
Simulating classification models to evaluate Predict-Then-Optimize methods
- 提出新算法模拟多分类器的预测输出
- 发现预测误差与最优解距离关系复杂
- 适合研究机器学习驱动决策系统的学者
优化中的不确定性常以随机参数形式建模。在预测-再优化框架中,机器学习模型的预测结果被用作这些参数的取值,从而将随机优化问题转化为确定性问题。该两阶段方法基于假设:预测越准确,得到的解越接近真实最优解。然而,在复杂约束优化场景下验证这一假设极具挑战,且常被忽视。模拟机器学习模型的预测结果,可在不实际训练模型的前提下,实验性分析预测误差对解质量的影响。本文在已有二分类模拟算法基础上,提出一种新的多分类器预测模拟算法,并通过计算实验评估其性能。结果显示,分类器性能可被合理模拟,尽管存在一定程度波动。进一步将该方法应用于机器调度问题的预测-再优化算法评估,发现预测误差与解距最优解的距离之间关系非线性,凸显了基于机器学习预测的决策系统设计与评估中的关键考量。
原文摘要 · Abstract (English)
Uncertainty in optimization is often represented as stochastic parameters in the optimization model. In Predict-Then-Optimize approaches, predictions of a machine learning model are used as values for such parameters, effectively transforming the stochastic optimization problem into a deterministic one. This two-stage framework is built on the assumption that more accurate predictions result in solutions that are closer to the actual optimal solution. However, providing evidence for this assumption in the context of complex, constrained optimization problems is challenging and often overlooked in the literature. Simulating predictions of machine learning models offers a way to (experimentally) analyze how prediction error impacts solution quality without the need to train real models. Complementing an algorithm from the literature for simulating binary classification, we introduce a new algorithm for simulating predictions of multiclass classifiers. We conduct a computational study to evaluate the performance of these algorithms, and show that classifier performance can be simulated with reasonable accuracy, although some variability is observed. Additionally, we apply these algorithms to assess the performance of a Predict-Then-Optimize algorithm for a machine scheduling problem. The experiments demonstrate that the relationship between prediction error and how close solutions are to the actual optimum is non-trivial, highlighting important considerations for the design and evaluation of decision-making systems based on machine learning predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。