arXiv:2505.17439cs.LGcs.AI2025-05
用强化学习动态优化救灾供应链,兼顾效率与公平性
Designing an efficient and equitable humanitarian supply chain dynamically via reinforcement learning
- 采用PPO强化学习算法实现动态调度
- 模型始终优先保障平均满意度
- 相比启发式算法更均衡高效
本研究利用强化学习中的PPO算法,动态设计高效且公平的人道主义供应链,并与启发式算法进行对比。结果表明,该模型在所有情境下均以提升平均满意度为首要目标,能有效平衡资源分配的公平性与整体效率,在应急响应中具有显著优势。
原文摘要 · Abstract (English)
This study designs an efficient and equitable humanitarian supply chain dynamically by using reinforcement learning, PPO, and compared with heuristic algorithms. This study demonstrates the model of PPO always treats average satisfaction rate as the priority.
强化学习救灾供应链动态调度
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。