arXiv:2607.18573cs.AI2026-07中稿 · presentation at AI…

机器学习在供应链延迟预测中是否优于按价值排序?

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

  • 用预测延迟严重度乘以价值进行排序,比仅按严重度排序更优
  • 在10%审查预算下,不同数据集表现差异大,最高提升10.1个百分点
  • 强调需先验证模型可学习性和校准性,再决定是否部署

延迟风险模型通常以预测准确率评估,但实际管理中更关注:当只能审查少数货件时,应优先检查哪些?本文在三个真实供应链场景(SCMS采购、DataCo物流、Olist电商)中,采用防泄露的滚动起源评估和1000次配对自助法置信区间,检验机器学习能否超越一个严格基准——仅按货件价值排序。结果表明,将预测延迟严重度与已知价值相乘(M1)的排序方式,在所有数据集中均优于仅按严重度排序,但未普遍优于单纯按价值排序。在10%审查预算下,M1相比仅按价值排序的增益为:SCMS -5.5个百分点,DataCo +10.1个百分点,Olist -4.9个百分点。这一差异与延迟严重度的可学习性一致:DataCo的R²为0.27,校准偏差+0.01天;而SCMS和Olist的R²约-0.02,校准偏差为负。嵌套交叉验证的成本敏感重训也未带来稳定改进。本文不提出新算法,而是提供一种部署前诊断与评估协议:价值排序应作为永久基准,机器学习必须通过可学习性与校准性的审计,并在防泄露滚动起源评估中达标后才能部署。

原文摘要 · Abstract (English)

Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones should a manager check first? We evaluate whether machine learning clears a demanding no-model baseline: inspect the highest-value shipments first. Across three real supply-chain contexts: SCMS procurement, DataCo logistics, and Olist e-commerce, we use leakage-controlled rolling-origin evaluation and 1000-sample paired bootstrap confidence intervals. Ranking by predicted delay severity times known value (M1) beats severity-only ranking in all three datasets, yet it does not generally beat value sorting. At a 10% review budget, M1 minus VALUE_ONLY is -5.5 percentage points (pp) for SCMS, +10.1 pp for DataCo, and -4.9 pp for Olist. The divide is consistent with severity learnability: DataCo has R^2 = 0.27 and calibration bias of +0.01 days, whereas SCMS and Olist have R^2 of approximately -0.02 and negative calibration bias. Nested-CV cost-sensitive retraining does not deliver a stable improvement over M1. Rather than proposing a new learning algorithm, this paper presents a deployment diagnostic and evaluation protocol. Value sorting should remain a permanent benchmark, and ML should be deployed only after severity learnability and calibration have been audited and the model clears that gate under leakage-controlled rolling-origin evaluation.

供应链机器学习排序优化评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。