用强化学习动态分配资源,兼顾情境约束与效率。
Situational-Constrained Sequential Resources Allocation via Reinforcement Learning
- 将情境约束建模为逻辑蕴含,动态惩罚违规行为。
- 在疫情医疗和农业农药分配中,约束满足率显著提升。
- 适合需要实时适应环境变化的资源调度场景。
序列资源分配面临情境约束的挑战,资源需求与优先级随上下文变化。本文提出新框架SCRL,将情境约束形式化为逻辑蕴含,并设计动态惩罚机制以减少违反约束的情况。为克服传统约束强化学习的局限,引入概率选择机制。在疫情期间医疗资源分配与农业农药分发两个场景中进行评估,实验表明SCRL在保持高资源效率的同时,显著优于现有基线方法,更有效满足约束条件,展现出在真实世界情境敏感决策任务中的应用潜力。
原文摘要 · Abstract (English)
Sequential Resource Allocation with situational constraints presents a significant challenge in real-world applications, where resource demands and priorities are context-dependent. This paper introduces a novel framework, SCRL, to address this problem. We formalize situational constraints as logic implications and develop a new algorithm that dynamically penalizes constraint violations. To handle situational constraints effectively, we propose a probabilistic selection mechanism to overcome limitations of traditional constraint reinforcement learning (CRL) approaches. We evaluate SCRL across two scenarios: medical resource allocation during a pandemic and pesticide distribution in agriculture. Experiments demonstrate that SCRL outperforms existing baselines in satisfying constraints while maintaining high resource efficiency, showcasing its potential for real-world, context-sensitive decision-making tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。