arXiv:2502.09724cs.LG2025-02ICML被引 8

为多目标强化学习设计可覆盖所有公平性权衡的策略组合,助决策者灵活选择最优方案。

Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning

  • 构建跨参数p的近似策略组合,统一处理不同公平性准则下的最优策略
  • 在合成与真实数据集上验证,组合能有效覆盖不同社会福利函数下的策略分布
  • 提供可调的近似精度、组合规模与计算效率的理论权衡,适合需灵活决策的场景

在强化学习的实际应用中,部署策略对不同利益相关方的影响各异,难以达成偏好聚合共识。广义p-均值是一类广泛使用的社会福利函数,适用于公平资源分配、人工智能对齐和决策制定,涵盖平等主义、纳什和功利主义等经典形式。然而,由于最优策略的结构与结果对p值高度敏感,决策者选择合适的福利函数极具挑战。为此,本文研究了α近似策略组合的概念——一组在p∈[−∞,1]范围内对所有广义p-均值近似最优的策略。我们提出计算此类组合的算法,并给出近似因子、组合大小与计算效率之间的理论保证。在合成与真实世界数据集上的实验表明,该方法能有效总结不同p值下的策略空间,使决策者更高效地导航这一复杂决策界面。

原文摘要 · Abstract (English)

In many real-world applications of reinforcement learning (RL), deployed policies have varied impacts on different stakeholders, creating challenges in reaching consensus on how to effectively aggregate their preferences. Generalized $p$-means form a widely used class of social welfare functions for this purpose, with broad applications in fair resource allocation, AI alignment, and decision-making. This class includes well-known welfare functions such as Egalitarian, Nash, and Utilitarian welfare. However, selecting the appropriate social welfare function is challenging for decision-makers, as the structure and outcomes of optimal policies can be highly sensitive to the choice of $p$. To address this challenge, we study the concept of an $α$-approximate portfolio in RL, a set of policies that are approximately optimal across the family of generalized $p$-means for all $p \in [-\infty, 1]$. We propose algorithms to compute such portfolios and provide theoretical guarantees on the trade-offs among approximation factor, portfolio size, and computational efficiency. Experimental results on synthetic and real-world datasets demonstrate the effectiveness of our approach in summarizing the policy space induced by varying $p$ values, empowering decision-makers to navigate this landscape more effectively.

强化学习多目标决策社会福利策略组合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。