用洛伦兹支配提升多目标强化学习的公平性与可扩展性
Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz Dominance
- 基于洛伦兹支配定义公平策略,支持灵活的公平偏好设置
- 在西安和阿姆斯特丹真实交通场景中验证,显著提升高维目标空间的效率
- 适用于需兼顾多方利益的复杂决策系统,如智慧城市调度
多目标强化学习(MORL)旨在学习一组在多个冲突目标间权衡的策略,但其计算复杂度随目标数量增加而急剧上升。当目标涉及个体或群体偏好时,公平性成为关键需求。本文提出一种融合公平性的原则性算法,利用洛伦兹支配识别奖励分配均衡的策略,并引入lambda-洛伦兹支配以支持灵活的公平偏好。我们发布了一个大规模真实世界交通规划环境,实验表明该方法能有效发现公平策略,在西安和阿姆斯特丹两个大城市中展现出优异的可扩展性。相比常见多目标方法,本方法在高维目标空间中表现更优。
原文摘要 · Abstract (English)
Multi-Objective Reinforcement Learning (MORL) aims to learn a set of policies that optimize trade-offs between multiple, often conflicting objectives. MORL is computationally more complex than single-objective RL, particularly as the number of objectives increases. Additionally, when objectives involve the preferences of agents or groups, incorporating fairness becomes both important and socially desirable. This paper introduces a principled algorithm that incorporates fairness into MORL while improving scalability to many-objective problems. We propose using Lorenz dominance to identify policies with equitable reward distributions and introduce lambda-Lorenz dominance to enable flexible fairness preferences. We release a new, large-scale real-world transport planning environment and demonstrate that our method encourages the discovery of fair policies, showing improved scalability in two large cities (Xi'an and Amsterdam). Our methods outperform common multi-objective approaches, particularly in high-dimensional objective spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。