用离线强化学习和微调大模型优化仓库排班,提升效率2.4%。
Learning to Staff: Offline Reinforcement Learning and Fine-Tuned LLMs for Warehouse Staffing Optimization
- 用历史数据训练基于Transformer的离线强化学习策略。
- 在模拟环境中实现比历史基准高2.4%的吞吐量提升。
- 大模型适合人类可读输入,可结合管理者偏好迭代优化。
我们研究机器学习方法以优化半自动化仓库分拣系统中的实时人员排班决策。在不同抽象层次上支持运营决策,各有权衡。我们评估两种方法,均在匹配的仿真环境中进行。第一种,使用离线强化学习在详细历史状态表示上训练自定义Transformer策略,在学习型模拟器中比历史基线提升2.4%吞吐量;在高流量仓库中,此改进带来显著节省。第二种,探索运行于抽象化、人类可读状态描述上的大语言模型(LLM),这类决策天然契合仓库经理基于高层运营摘要的判断。我们系统比较提示工程、自动提示优化及微调策略。仅靠提示效果不足,但监督微调结合模拟器生成偏好数据的直接偏好优化,使性能达到或略超历史基线。结果表明,两种路径均可实现人工智能辅助运营决策:离线强化学习擅长任务特定架构;大模型支持人类可读输入,并可与包含经理偏好的迭代反馈环结合。
原文摘要 · Abstract (English)
We investigate machine learning approaches for optimizing real-time staffing decisions in semi-automated warehouse sortation systems. Operational decision-making can be supported at different levels of abstraction, with different trade-offs. We evaluate two approaches, each in a matching simulation environment. First, we train custom Transformer-based policies using offline reinforcement learning on detailed historical state representations, achieving a 2.4% throughput improvement over historical baselines in learned simulators. In high-volume warehouse operations, improvements of this size translate to significant savings. Second, we explore LLMs operating on abstracted, human-readable state descriptions. These are a natural fit for decisions that warehouse managers make using high-level operational summaries. We systematically compare prompting techniques, automatic prompt optimization, and fine-tuning strategies. While prompting alone proves insufficient, supervised fine-tuning combined with Direct Preference Optimization on simulator-generated preferences achieves performance that matches or slightly exceeds historical baselines in a hand-crafted simulator. Our findings demonstrate that both approaches offer viable paths toward AI-assisted operational decision-making. Offline RL excels with task-specific architectures. LLMs support human-readable inputs and can be combined with an iterative feedback loop that can incorporate manager preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。