首个直接优化非线性公平目标的离线多目标强化学习框架
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
- 通过分布修正估计,联合优化公平性与分布正则化
- 无需偏好权重,实现稳定高效的离线学习
- 在多个基准上显著提升公平性表现
多目标强化学习(MORL)旨在处理存在冲突目标的场景,通常采用线性加权将多维回报转化为标量信号。然而,该方法难以捕捉如纳什社会福利或最大最小公平性等公平导向目标,这些目标需要非线性、非可加的权衡。尽管已有在线算法针对特定公平目标,但在固定数据集上优化非线性福利准则的统一离线方法仍属空白。本文提出FairDICE,首个直接优化非线性福利目标的离线MORL框架。它利用分布修正估计,同时兼顾福利最大化与分布正则化,实现无需显式偏好权重或遍历权重搜索的稳定且样本高效的学习。在多个离线基准测试中,FairDICE相较于现有基线展现出更强的公平性感知性能。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) aims to optimize policies in the presence of conflicting objectives, where linear scalarization is commonly used to reduce vector-valued returns into scalar signals. While effective for certain preferences, this approach cannot capture fairness-oriented goals such as Nash social welfare or max-min fairness, which require nonlinear and non-additive trade-offs. Although several online algorithms have been proposed for specific fairness objectives, a unified approach for optimizing nonlinear welfare criteria in the offline setting-where learning must proceed from a fixed dataset-remains unexplored. In this work, we present FairDICE, the first offline MORL framework that directly optimizes nonlinear welfare objective. FairDICE leverages distribution correction estimation to jointly account for welfare maximization and distributional regularization, enabling stable and sample-efficient learning without requiring explicit preference weights or exhaustive weight search. Across multiple offline benchmarks, FairDICE demonstrates strong fairness-aware performance compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。