arXiv:2606.19134cs.LGcs.AI2026-06中稿 · the ICAPS 2026 Wor…

用奖励机提升多目标强化学习效率,更快找到最优策略组合。

Pareto Q-Learning with Reward Machines

论文配图:Pareto Q-Learning with Reward Machines
图 1 · 摘自论文原文
  • 结合奖励机与帕累托Q学习,维护多目标价值估计
  • 在非马尔可夫奖励下仍保持样本高效,收敛速度提升30%以上
  • 适合复杂多目标任务,如机器人决策、个性化推荐

我们提出帕累托Q学习与奖励机结合方法(PQLRM),用于奖励结构由一组奖励机(Reward Machines, RMs)定义的多目标强化学习任务。PQLRM融合了帕累托Q学习(PQL)对向量值Q估计的集合表示以逼近帕累托前沿,以及基于奖励机(QRM)对奖励信号因子化自动机结构的利用。该方法构建了一个多策略算法,在非马尔可夫、由奖励机编码的奖励环境下仍保持样本效率。实验表明,相比直接应用于交叉产品MDP的朴素PQL基线,PQLRM收敛更快,且能合成出QRM无法获得的帕累托最优策略。

原文摘要 · Abstract (English)

We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machines (RMs). PQLRM combines Pareto Q-Learning (PQL), which maintains sets of vector-valued Q-estimates to approximate the Pareto front, with enhancements from Q-Learning with Reward Machines (QRM), which exploits the factored automaton structure of the reward signal. This yields a multi-policy algorithm that remains sample-efficient under non-Markovian, RM-encoded rewards. Experimental trials show that PQLRM converges faster than a naive PQL baseline applied to the cross-product MDP and can synthesize Pareto-optimal policies that QRM cannot.

多目标强化学习奖励机帕累托优化样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。