用滚动优化方法直接最小化业务流程周期时间
A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes
- 基于执行轨迹评估的迭代策略优化方法
- 在6个可计算最优解场景中学习到最优策略
- 适合复杂真实业务流程资源分配问题
资源分配在缩短业务流程周期时间和提升效率方面至关重要。近年来,深度强化学习(DRL)成为优化业务流程资源分配策略的强大工具。在DRL框架中,智能体通过与环境交互,仅依据奖励信号学习策略。然而,现有算法不适用于动态环境,且依赖人工设计的奖励函数,易导致奖励与目标错位,引发次优决策。为此,本文提出一种基于滚动的DRL算法和直接分解周期时间目标的奖励函数,使试错式奖励工程成为历史。该算法通过评估不同动作后的执行轨迹迭代优化策略。我们在六个可计算最优解的场景及一系列日益复杂的现实规模流程模型上验证了该方法。结果表明,该算法能在所有场景中学习到最优策略,并在真实业务流程上表现优于或匹配最佳启发式方法。
原文摘要 · Abstract (English)
Resource allocation plays a critical role in minimizing cycle time and improving the efficiency of business processes. Recently, Deep Reinforcement Learning (DRL) has emerged as a powerful technique to optimize resource allocation policies in business processes. In the DRL framework, an agent learns a policy through interaction with the environment, guided solely by reward signals that indicate the quality of its decisions. However, existing algorithms are not suitable for dynamic environments such as business processes. Furthermore, existing DRL-based methods rely on engineered reward functions that approximate the desired objective, but a misalignment between reward and objective can lead to undesired decisions or suboptimal policies. To address these issues, we propose a rollout-based DRL algorithm and a reward function to optimize the objective directly. Our algorithm iteratively improves the policy by evaluating execution trajectories following different actions. Our reward function directly decomposes the objective function of minimizing the cycle time, such that trial-and-error reward engineering becomes unnecessary. We evaluated our method in six scenarios, for which the optimal policy can be computed, and on a set of increasingly complex, realistically sized process models. The results show that our algorithm can learn the optimal policy for the scenarios and outperform or match the best heuristics on the realistically sized business processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。