用事件驱动强化学习优化芯片制造长时序控制,提升产线吞吐与利用率。
Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication

- 以事件驱动方式建模制造流程,设计可适配多种算法的时序差分框架。
- 在离线与在线训练下,产线吞吐和设备利用率均显著提升。
- 适用于复杂自适应系统控制,适合工业界长期决策场景。
强化学习有望优化大规模系统中的序列决策。半导体制造系统是随机性强、约束严苛的环境,异构晶圆需经过数百道工序穿越庞大的设备网络。这些特性导致复杂的高维决策问题,伴随反馈延迟和长时序需求,使生产规划与控制极具挑战。本文提出一种面向多目标策略优化的深度强化学习框架。具体而言,将控制建模为集中式智能体问题,由核心策略协调全局决策,系统演化则以离散事件驱动的互联时序过程表示。为此,开发了一种定制化的事件驱动时序差分方法,具备通用性,可与多种策略优化方法在不同训练设置下集成。我们测试了多种主流无模型算法在该框架中的表现,并通过高保真仿真评估其在多样化真实工业场景下的有效性。在广泛的验证实验中,离线与在线训练的智能体均在吞吐量和设备利用率上实现显著且一致的提升。进一步分析了不同训练阶段的性能与泛化能力,厘清了各类强化学习方法的相对优势。总体结果表明,该框架具备可扩展性、通用性与迁移性,适用于事件驱动的复杂自适应系统控制。
原文摘要 · Abstract (English)
Reinforcement learning promises to optimize sequential decisions in large-scale systems. Semiconductor manufacturing systems are stochastic and highly constrained environments where heterogeneous wafers traverse hundreds of processing steps across extensive equipment networks. These characteristics yield complex, high-dimensional decision problems with delayed feedback and long-horizon requirements, complicating production planning and control. We propose a deep reinforcement learning framework for multi-objective policy optimization at this scale. Specifically, we formulate control as a centralized-agent problem, where a core policy coordinates system-wide decisions, while system evolution is represented as an interconnected temporal process driven by discrete events. Accordingly, we develop a tailored event-driven temporal-difference formulation that remains general and can be integrated with various policy optimization methods under relevant training settings. We investigate several core model-free algorithms incorporated into this framework and evaluate their effectiveness using high-fidelity simulations of diverse, industry-real operating scenarios. Across extensive validation experiments, agents trained in both offline and online settings show significant and consistent gains in throughput and utilization. We further evaluate performance and generalization across training phases, clarifying the relative strengths of alternative reinforcement learning formulations and algorithms. Overall, the results support the scalability, generality, and transferability of the proposed framework for controlling event-driven complex adaptive systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。