arXiv:2504.02662cs.LGmath.OC2025-04被引 3

用人类经验指导强化学习,提升决策可信度与效果

Integrating Human Knowledge Through Action Masking in Reinforcement Learning for Operations Research

  • 通过动作屏蔽引入专家经验,引导智能体选择合理行动
  • 在三种运筹问题中,融合人类知识后性能显著提升
  • 过度限制动作会抑制探索,反而降低整体表现

强化学习在运筹学问题中具有强大潜力,但实际应用常因缺乏用户信任而受阻。本文探讨通过动作屏蔽融入人类专家知识的利弊。尽管动作屏蔽通常用于排除无效动作,其整合人类经验的能力仍待挖掘。人类知识常体现为启发式规则,能推荐特定情境下的合理近优动作。强制执行此类动作可增强人力对模型决策的信任,但严格限制可能阻碍智能体发现更优策略,导致性能下降。研究分析了喷漆车间调度、峰值负荷管理与库存管理三个不同特性的问题。结果表明,通过动作屏蔽融入人类知识可显著优于无屏蔽训练的策略;在动作受限场景(如某些动作次数有限)中,动作屏蔽对学习有效策略至关重要;同时提醒,过强的动作约束可能导致次优结果。

原文摘要 · Abstract (English)

Reinforcement learning (RL) provides a powerful method to address problems in operations research. However, its real-world application often fails due to a lack of user acceptance and trust. A possible remedy is to provide managers with the possibility of altering the RL policy by incorporating human expert knowledge. In this study, we analyze the benefits and caveats of including human knowledge via action masking. While action masking has so far been used to exclude invalid actions, its ability to integrate human expertise remains underexplored. Human knowledge is often encapsulated in heuristics, which suggest reasonable, near-optimal actions in certain situations. Enforcing such actions should hence increase trust among the human workforce to rely on the model's decisions. Yet, a strict enforcement of heuristic actions may also restrict the policy from exploring superior actions, thereby leading to overall lower performance. We analyze the effects of action masking based on three problems with different characteristics, namely, paint shop scheduling, peak load management, and inventory management. Our findings demonstrate that incorporating human knowledge through action masking can achieve substantial improvements over policies trained without action masking. In addition, we find that action masking is crucial for learning effective policies in constrained action spaces, where certain actions can only be performed a limited number of times. Finally, we highlight the potential for suboptimal outcomes when action masks are overly restrictive.

强化学习运筹优化人机协同动作屏蔽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。