预测排行榜冠军未必适合部署,新协议确保决策可靠
From Forecasting Leaderboards to Deployment Decisions: A Fail-Closed Certification Protocol

- 设计保守认证流程,判断预测模型能否直接用于部署
- 实测发现155次预测领先者在部署时反而表现更差
- 适合关注模型落地可靠性的工程师和决策者
预测排行榜按预测质量排序模型,但其胜者常被误认为可直接部署的最优选择。当预测结果通过固定决策接口(如告警阈值、前k名预算或切换成本策略)时,这种解读可能失效。本文研究何种条件下预测优胜者可被认证为部署可用的首选。提出一种‘失败封闭’认证协议,其门限基于充分证据:摩擦导致的非平局、统计支持、可复现的部署侧反转。以Traffic-Hourly为例,零摩擦下模型排名一致,但存在切换成本时,预测胜者反而部署表现更差。锁定原生审计测试过度宣称:在22个验证候选与362个完整网格单元中,155次看似预测/部署胜者反转被认证前阻断。贡献不在于新模型或度量,而是一种保守的决策协议,用于判断预测排行榜胜者是否可作部署首推。
原文摘要 · Abstract (English)
Forecasting leaderboards rank models by predictive quality, but their winners are often read as deployment-ready top-1 advice. That reading can fail when forecasts are passed through a fixed decision interface, such as an alert threshold, a top-k budget, or a switching-cost policy. We study when a forecast-side winner can be certified as deployment-actionable for a specified interface and deployed utility. We introduce a fail-closed certification protocol whose gates are sufficient evidential conditions for a strong claim: a friction-caused, non-tie, statistically supported, and recurrent deployment-side reversal. Traffic-Hourly provides a certified anchor: winners agree at zero friction, but positive switching friction makes the forecast winner deployed-suboptimal. A locked native audit tests overclaiming: across 22 verified candidates and 362 full-grid cells, 155 apparent forecast/deployment winner inversions are blocked before certification. The contribution is not a new forecaster, metric, or universal utility, but a conservative protocol for deciding when forecasting leaderboard winners should be read as deployment-actionable top-1 advice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。