arXiv:2512.19738cs.LG2025-12

用强化学习动态调缓冲区,提升快递站包裹量预测准确率

OpComm: A Reinforcement Learning Framework for Adaptive Buffer Control in Warehouse Volume Forecasting

  • 结合预测模型与强化学习,自动调节站点缓冲区
  • WAPE降低21.65%,减少缺货事件
  • 生成可解释报告,适合物流管理者决策

精准预测配送站包裹量对末端物流至关重要,误差会导致资源浪费、成本上升和延误。本文提出OpComm框架,融合监督学习与基于强化学习的缓冲区控制,以及生成式AI通信模块。采用LightGBM回归模型生成站点级需求预测,作为近端策略优化(PPO)代理的上下文,从离散动作集中选择缓冲区水平。奖励函数更严厉惩罚欠缓冲,反映实际中未满足需求与资源低效之间的权衡。通过蒙特卡洛更新机制反馈站点结果,实现策略持续优化。为增强可解释性,生成式AI层基于SHAP特征归因生成管理层摘要与情景分析。在400多个站点上,相比人工预测,OpComm将加权绝对百分比误差(WAPE)降低21.65%,减少欠缓冲事件,并提升决策透明度。本研究展示了情境化强化学习与预测建模结合,在高风险物流环境中弥合统计严谨性与实际决策能力的有效路径。

原文摘要 · Abstract (English)

Accurate forecasting of package volumes at delivery stations is critical for last-mile logistics, where errors lead to inefficient resource allocation, higher costs, and delivery delays. We propose OpComm, a forecasting and decision-support framework that combines supervised learning with reinforcement learning-based buffer control and a generative AI-driven communication module. A LightGBM regression model generates station-level demand forecasts, which serve as context for a Proximal Policy Optimization (PPO) agent that selects buffer levels from a discrete action set. The reward function penalizes under-buffering more heavily than over-buffering, reflecting real-world trade-offs between unmet demand risks and resource inefficiency. Station outcomes are fed back through a Monte Carlo update mechanism, enabling continual policy adaptation. To enhance interpretability, a generative AI layer produces executive-level summaries and scenario analyses grounded in SHAP-based feature attributions. Across 400+ stations, OpComm reduced Weighted Absolute Percentage Error (WAPE) by 21.65% compared to manual forecasts, while lowering under-buffering incidents and improving transparency for decision-makers. This work shows how contextual reinforcement learning, coupled with predictive modeling, can address operational forecasting challenges and bridge statistical rigor with practical decision-making in high-stakes logistics environments.

强化学习物流预测缓冲控制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。