用进化算法优化多目标供应链,提升解的多样性和适应性。
Meta-Reinforcement Learning via Evolution for Multi-Objective Combinatorial Supply Chain Optimisation
- 通过种群进化优化权重向量,生成多个元策略协同求解。
- 在复杂场景下提升超体积32%,帕累托前沿分布更优。
- 适合需兼顾经济、环境与社会目标的供应链优化问题。
元强化学习在多目标优化中具有快速适应环境变化和偏好设置的潜力。然而,传统少样本方法通常从单一共享的元策略微调,限制了解的多样性,难以充分探索帕累托前沿,尤其在高维组合优化问题如供应链优化中表现不佳。本文提出一种基于种群的元强化学习框架,结合分解与标量化权重空间中的进化搜索。该框架维护一组权重向量,每个对应一个通过梯度元学习训练的独立元策略,并通过精英选择、交叉和变异迭代优化,依据超体积和熵贡献进行指导。我们在包含经济、环境与社会冲突目标的多目标供应链场景中评估该方法,并进一步测试其在标准强化学习任务上的泛化能力。结果表明,所提方法获得更丰富、分布更优的帕累托前沿近似,显著提升跨任务适应能力,在复杂情况下相比元多目标强化学习提升超体积达32%,且所有对比方法中平均豪斯多夫距离最低。
原文摘要 · Abstract (English)
Meta-reinforcement learning is a promising approach to multi-objective optimisation because it enables rapid policy adaptation across changing environments and preference settings. However, conventional few-shot methods usually fine-tune from a single shared meta-policy, which can reduce solution diversity and limit exploration of the Pareto front, especially in high-dimensional combinatorial problems such as supply chain optimisation. We propose a population-based Meta-reinforcement learning framework that combines decomposition with evolutionary search in scalarisation weight space. The framework maintains a population of weight vectors, each associated with a distinct meta-policy trained through gradient-based meta-learning, and iteratively refines this population through elitist selection, crossover, and mutation guided by hypervolume and entropy contributions. We evaluate the method in a multi-objective supply chain setting with conflicting economic, environmental, and social goals, and further test its generality on standard reinforcement learning problems. The results show that the proposed approach yields more diverse, better distributed Pareto front approximations, improves cross-task adaptation, increases hypervolume by up to 32\% over Meta-multi-objective reinforcement learning in the complex case, and attains the lowest average Hausdorff distance among all compared methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。