对比了两种优化建模方法在模型不准确时的表现,发现新方法更稳健。
Dissecting the Impact of Model Misspecification in Data-driven Optimization
- 用决策误差直接优化,而非传统估计误差
- 模型偏差大时,新方法可降低两倍关键损失项
- 适合模型不确定的现实场景,如供应链、金融决策
数据驱动优化通过将机器学习模型转化为决策,优化预估成本。传统方法使用最大似然拟合分布模型,再代入优化问题;近年出现的新方法则整合估计与优化过程,直接最小化决策误差。尽管直觉上合理,其统计优势尚不明确。本文通过有限样本的尾部后悔界分析,利用高阶展开和最新的Berry-Esseen定理,揭示当底层模型存在偏差时,集成方法在后悔的前两个主导项上具有“普遍双益”;而模型几乎正确时,传统方法仍可能更优。该比较为机器学习在决策中的应用提供了理论指导。
原文摘要 · Abstract (English)
Data-driven optimization aims to translate a machine learning model into decision-making by optimizing decisions on estimated costs. Such a pipeline can be conducted by fitting a distributional model which is then plugged into the target optimization problem. While this fitting can utilize traditional methods such as maximum likelihood, a more recent approach uses estimation-optimization integration that minimizes decision error instead of estimation error. Although intuitive, the statistical benefit of the latter approach is not well understood yet is important to guide the prescriptive usage of machine learning. In this paper, we dissect the performance comparisons between these approaches in terms of the amount of model misspecification. In particular, we show how the integrated approach offers a ``universal double benefit'' on the top two dominating terms of regret when the underlying model is misspecified, while the traditional approach can be advantageous when the model is nearly well-specified. Our comparison is powered by finite-sample tail regret bounds that are derived via new higher-order expansions of regrets and the leveraging of a recent Berry-Esseen theorem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。