arXiv:2603.27066cs.LGcs.AI2026-03被引 9

用深度强化学习动态匹配制造资源,提升产能利用率。

Dynamic resource matching in manufacturing using deep reinforcement learning

  • 基于深度强化学习构建多期多对多资源匹配模型。
  • 新算法在大小规模实验中均获更高收益和更快收敛。
  • 适合智能制造、资源调度场景的从业者参考。

匹配在多个行业的资源逻辑分配中具有重要意义,尤其在制造领域,产能共享备受关注。本文研究制造资源的需求-产能类型动态匹配问题,将其建模为多周期、多对多的序贯决策过程。该问题状态空间与动作空间庞大,难以准确建模各类需求的联合分布。针对维度灾难及转移动态难以显式建模的问题,采用无模型深度强化学习方法求解最优匹配策略。为解决Q-learning中不可行动作及初始偏差导致的收敛慢问题,引入基于先验策略的领域知识惩罚与符合供需约束的不可行性惩罚。理论上证明了改进Q-learning在小规模问题上的收敛性与性能保障;对于大规模问题,将该方法融入深度确定性策略梯度(DDPG)算法,提出领域知识引导的DDPG(DKDDPG)。计算实验包括小规模与大规模测试,结果表明DKDDPG持续优于传统DDPG及其他强化学习算法,在奖励、时间与训练轮次上均表现更优。

原文摘要 · Abstract (English)

Matching plays an important role in the logical allocation of resources across a wide range of industries. The benefits of matching have been increasingly recognized in manufacturing industries. In particular, capacity sharing has received much attention recently. In this paper, we consider the problem of dynamically matching demand-capacity types of manufacturing resources. We formulate the multi-period, many-to-many manufacturing resource-matching problem as a sequential decision process. The formulated manufacturing resource-matching problem involves large state and action spaces, and it is not practical to accurately model the joint distribution of various types of demands. To address the curse of dimensionality and the difficulty of explicitly modeling the transition dynamics, we use a model-free deep reinforcement learning approach to find optimal matching policies. Moreover, to tackle the issue of infeasible actions and slow convergence due to initial biased estimates caused by the maximum operator in Q-learning, we introduce two penalties to the traditional Q-learning algorithm: a domain knowledge-based penalty based on a prior policy and an infeasibility penalty that conforms to the demand-supply constraints. We establish theoretical results on the convergence of our domain knowledge-informed Q-learning providing performance guarantee for small-size problems. For large-size problems, we further inject our modified approach into the deep deterministic policy gradient (DDPG) algorithm, which we refer to as domain knowledge-informed DDPG (DKDDPG). In our computational study, including small- and large-scale experiments, DKDDPG consistently outperformed traditional DDPG and other RL algorithms, yielding higher rewards and demonstrating greater efficiency in time and episodes.

制造资源强化学习动态匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。