MIRACL让供应链优化模型快速适应新任务,少样本学习效果更好。
MIRACL: A Diverse Meta-Reinforcement Learning for Multi-Objective Multi-Echelon Combinatorial Supply Chain Optimisation
- 分层分解任务,用元学习实现跨任务快速适应
- 在测试中比传统方法高10%的超体积指标,提升5%预期收益
- 适合需要快速响应多目标变化的动态决策场景
多目标强化学习(MORL)在多层级组合供应链优化中表现优异,但面对动态环境时需针对每个任务重新训练,且计算成本高。本文提出MIRACL(Meta multI-objective Reinforcement leArning with Composite Learning),一种分层式元多目标强化学习框架,支持在多样化任务间实现少样本泛化。该框架将每项任务分解为结构化子问题,通过基于帕累托的自适应策略,在元训练与微调中促进多样性,从而高效适配全局策略。据我们所知,这是首个将元多目标强化学习与此类机制结合于组合优化的研究。尽管在供应链领域验证,理论框架可适用于更广泛的动态多目标决策问题。实验表明,相较于传统MORL基线,MIRACL在简单到中等复杂度任务中表现更优,超体积指标最高提升10%,预期效用改善5%。结果凸显了其在多目标问题中实现鲁棒、高效适应的潜力。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) is effective for multi-echelon combinatorial supply chain optimisation, where tasks involve high dimensionality, uncertainty, and competing objectives. However, its deployment in dynamic environments is hindered by the need for task-specific retraining and substantial computational cost. We introduce MIRACL (Meta multI-objective Reinforcement leArning with Composite Learning), a hierarchical Meta-MORL framework that allows for a few-shot generalisation across diverse tasks. MIRACL decomposes each task into structured subproblems for efficient policy adaptation and meta-learns a global policy across tasks using a Pareto-based adaptation strategy to encourage diversity in meta-training and fine-tuning. To our knowledge, this is the first integration of Meta-MORL with such mechanisms in combinatorial optimisation. Although validated in the supply chain domain, MIRACL is theoretically domain-agnostic and applicable to broader dynamic multi-objective decision-making problems. Empirical evaluations show that MIRACL outperforms conventional MORL baselines in simple to moderate tasks, achieving up to 10% higher hypervolume and 5% better expected utility. These results underscore the potential of MIRACL for robust, efficient adaptation in multi-objective problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。