统一非线性多目标强化学习框架,解决复杂权衡优化难题
AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL

- 提出AET三步分解框架,统一SER与ESR两类目标
- 基于密度比估计实现离线数据上的可训练算法
- 适合需处理风险、公平等复杂权衡的决策场景
在多目标强化学习中优化非线性偏好对捕捉风险规避或公平性等复杂权衡至关重要。然而,非线性特性长期导致非线性多目标强化学习(MORL)分为两类范式:标量期望回报(SER)和期望标量回报(ESR)。SER需要全局优化,而ESR要求非马尔可夫策略,造成优化策略碎片化。本文通过聚合-期望-变换(AET)框架,将两类标准通过标量化的三分解统一,为一般非线性MORL提供理论基础。在此基础上,提出可训练的离线强化学习算法AETDICE,利用增强状态空间中的DICE式密度比估计,实现从静态数据集进行样本化优化。该框架突破长期存在的技术壁垒,有效捕捉了由AET框架引发的各类权衡,是现有方法无法实现的。
原文摘要 · Abstract (English)
Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk aversion or fairness. However, such non-linearity has historically bifurcated nonlinear MORL objectives into two distinct paradigms: Scalarized Expected Return (SER) and Expected Scalarized Return (ESR). While SER requires global-level optimization and ESR requires non-Markovian policies, leading to fragmented optimization strategies, we bridge this divide through the Aggregation-Expectation-Transformation (AET) framework. By unifying both criteria through a tripartite decomposition of scalarization, AET provides a principled foundation for general nonlinear MORL. Building on this framework, we propose AETDICE, a tractable offline RL algorithm for AET objectives. By utilizing DICE-style density-ratio estimation in an augmented state space, AETDICE enables sample-based optimization from static datasets. Our framework resolves long-standing barriers and captures respective trade-offs induced by AET framework, which existing methods fail to address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。