arXiv:2605.31388cs.LG2026-05中稿 · ICML

提出兼顾公平与约束的多目标强化学习框架,解决冲突目标间的权衡问题。

Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion

  • 引入最大最小准则,优化多个冲突目标间的公平性
  • 理论证明算法收敛,在表格环境中验证有效性
  • 适用于建筑温控、运动控制等需满足约束的实际场景

多目标强化学习(MORL)通过同时优化多个常有冲突的目标扩展了标准强化学习。尽管最大最小MORL已成为提升公平性的有效方法,但其在需显式满足约束条件时应用受限。本文提出一种融合最大最小准则与显式约束满足的MORL框架,建立了该框架的理论基础,并通过收敛性分析及表格环境中的实验加以验证。进一步在模拟建筑热控、多目标运动控制和碳排放感知交通管理中展示了该方法的实用性。在这些场景中,该方法有效平衡了公平性与约束满足,在多目标决策中表现优异。

原文摘要 · Abstract (English)

Multi-Objective Reinforcement Learning (MORL) extends standard RL by optimizing policies with respect to multiple, often conflicting, objectives. While max-min MORL has emerged as an effective approach for promoting fairness, its applicability remains limited, particularly when constraints must be incorporated. In this paper, we propose a MORL framework that integrates the max-min criterion with explicit constraint satisfaction. We establish a theoretical foundation for the proposed framework and validate the resulting algorithm through convergence analysis and experiments in tabular settings. We further demonstrate the practical relevance of our approach in simulated building thermal control, multi-objective locomotion control, and greenhouse-gas-emission-aware traffic management. Across these domains, our method effectively balances fairness and constraint satisfaction in multi-objective decision-making.

强化学习多目标约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。