arXiv:2510.14837cs.LGcs.AI2025-10AAAI被引 40

提出可处理噪声奖励的随机奖励机器,提升复杂任务学习效率。

Reinforcement Learning with Stochastic Reward Machines

  • 用约束求解法学习最小化随机奖励机器,适应噪声环境
  • 在两个案例中表现优于现有方法和朴素处理方式
  • 适合需要长期依赖奖励的强化学习场景

奖励机器是解决奖励稀疏且依赖复杂动作序列的强化学习问题的成熟工具。然而,现有学习算法假设奖励完全无噪声,这一理想化设定限制了实际应用。为此,本文提出一种新型奖励机器——随机奖励机器,并设计了一种基于约束求解的学习算法。该算法从强化学习智能体的探索行为中学习最小化的随机奖励机器,可与现有强化学习算法无缝结合,在极限情况下保证收敛到最优策略。我们在两个案例研究中验证了该算法的有效性,结果表明其性能显著优于现有方法及处理噪声奖励的朴素策略。

原文摘要 · Abstract (English)

Reward machines are an established tool for dealing with reinforcement learning problems in which rewards are sparse and depend on complex sequences of actions. However, existing algorithms for learning reward machines assume an overly idealized setting where rewards have to be free of noise. To overcome this practical limitation, we introduce a novel type of reward machines, called stochastic reward machines, and an algorithm for learning them. Our algorithm, based on constraint solving, learns minimal stochastic reward machines from the explorations of a reinforcement learning agent. This algorithm can easily be paired with existing reinforcement learning algorithms for reward machines and guarantees to converge to an optimal policy in the limit. We demonstrate the effectiveness of our algorithm in two case studies and show that it outperforms both existing methods and a naive approach for handling noisy reward functions.

强化学习奖励机器噪声处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。