让调度模型在异常时自动参考规则、回放或AI建议,更快恢复生产。
Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery

- 在推理阶段通过修正动作概率,动态注入外部恢复指导。
- 使用自定义装配线环境测试,异常恢复时间减少23%,准时交付率提升18%。
- 支持多种知识源协同,适合工业调度系统快速集成外部经验。
工业装配线在设备故障、人员缺勤和紧急订单下需及时决策。现有方法或依赖僵化的手工恢复逻辑,或学习自适应策略但无法在决策时有效利用异构外部恢复知识,导致异常恢复时间(ART)长、准时交付率(OTD)低。为此,我们提出一种阶段感知的引导注入框架,通过在评估阶段的对数几率层对训练好的循环MAPPO(RMAPPO)调度策略进行动作偏差调整。该框架为基于规则、回放数据和在线LLM的引导提供统一的决策时接口,并仅在异常与恢复阶段激活干预。在自定义AssemblyLineEnv上的实验表明,高质量规则引导带来最大收益,回放引导在可用性不足时平滑退化,而在线LLM引导仍能提供显著的中间改进。结果证明,决策时引导注入可在不重训练智能体的情况下有效利用异构恢复提示。
原文摘要 · Abstract (English)
Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders. Existing methods either rely on rigid handcrafted recovery logic or learn adaptive policies that do not readily exploit heterogeneous external recovery knowledge at decision time to reduce abnormal recovery time (ART) and preserve on-time delivery (OTD). To address this gap, we propose a phase-aware guidance injection framework that augments a trained recurrent MAPPO (RMAPPO) scheduling policy through logit-level action bias during evaluation. The framework provides a unified decision-time interface for rule-based, replay-based, and online LLM-based guidance, while activating intervention only during abnormal and recovery phases. Experiments on a custom AssemblyLineEnv show that high-quality rule guidance yields the strongest gains, replay-based guidance degrades smoothly under imperfect availability, and online LLM guidance still provides useful intermediate improvements. These results show that decision-time guidance injection can exploit heterogeneous recovery hints without redesigning the actor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。