提出词典序奖励框架,解决多目标强化学习中单一奖励失效问题。
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
- 基于词典序偏好构建多维奖励函数,无需连续性假设。
- 证明二维奖励可表征不可用标量奖励表示的偏好关系。
- 适用于多目标决策场景,尤其适合对目标有优先级的智能体。
近期研究通过期望效用理论将奖励视为效用,形式化了奖励假设。Hausner 的奠基性工作表明,放弃连续性公理会导致效用为任意维度的词典序向量的推广。本文进一步识别出一个简单且实用的条件:当偏好无法用标量奖励表示时,必须采用二维奖励函数。我们对在偏好无记忆性假设下的马尔可夫决策过程(MDPs)中此类奖励函数进行了完整刻画,并扩展至 d 维情形。此外,我们证明该设定下的最优策略保留了许多标量奖励情形的优良性质,而约束型马尔可夫决策过程(CMDP)中的策略则不具备这些特性。
原文摘要 · Abstract (English)
Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。