用少量数据实时修正预测,提升零售需求预报精度与库存效率。
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting

- 先预测后修正,结合少样本上下文强化与掩码更新策略。
- 在多种需求模式下降低9.52%均方根误差,优于基线模型。
- 适合需要快速响应的零售场景,尤其适用于标签稀疏初期。
当需求变化速度超过静态模型重训练能力时,零售需求预测依然困难,尤其是在需求初期新标签稀疏的情况下。为解决此问题,本文提出一种预测-修正(PtC)框架,保留第一阶段机器学习预测结果,并采用少样本连续上下文博弈修正策略,结合相似商品增强与顶-p掩码更新机制。在沃尔玛零售数据及独家饮料数据集上,PtC在稳定高量、稳定低量和波动间歇需求模式下均显著降低MAPE、MAE和RMSE;消融实验显示,平均RMSE较纯机器学习基线改善9.52%;在测试的提前期条件下,其库存成本低于基础库存、近端策略优化和软演员-评论家策略。结果表明,在无需完全重训基模型的前提下,在线修正可弥合离线学习与实时决策间的差距。
原文摘要 · Abstract (English)
Retail demand forecasting remains difficult when demand shifts faster than static forecasting models can be retrained, especially in early demand cycles where newly observed labels are sparse. To address this, this study aims to improve adaptive retail forecasting by proposing a predict-then-correct (PtC) framework that retains a first-stage machine learning (ML) forecast and applies a few-shot continuous contextual bandit correction policy with similar-SKUs augmentation and top-p masked updating. Across Walmart retail data and an exclusive beverage dataset, PtC delivers statistically significant reductions in MAPE, MAE, and RMSE across stable & high volume, stable & low volume, and erratic & intermittent demand patterns, improves average RMSE by 9.52% over the ML-only baseline in the ablation study, and yields lower inventory costs than base-stock, proximal policy optimization, and soft actor-critic policies under the tested lead-time settings. These findings show that online forecast correction can bridge offline demand learning and real-time retail decision-making by adapting to sparse feedback without fully retraining the base forecasting model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。