用物理规律增强奖励机制,让强化学习更高效。
Physics-Informed Reward Machines
- 基于物理规律设计符号化奖励机,提升目标表达能力。
- 结合反事实经验与奖励塑形,加速训练阶段的奖励获取。
- 适用于需高效学习的物理控制任务,尤其适合机器人等场景。
奖励机器(RMs)为强化学习中的非马尔可夫奖励提供了结构化表达方式,提升了表达力和可编程性。它将环境已知信息(通过奖励机制体现)与未知部分(需采样探索)分离,支持反事实经验生成和奖励塑形等技术,降低样本复杂度并加快学习速度。本文提出物理信息奖励机器(pRM),一种符号化机器,用于表达复杂学习目标与奖励结构,使强化学习更具可编程性、表达力和效率。我们设计了可利用pRM进行反事实经验生成与奖励塑形的强化学习算法。实验表明,这些方法显著加速了强化学习训练阶段的奖励获取。在有限与连续物理环境中验证了pRM的表达力与有效性,证明其能显著提升多个控制任务的学习效率。
原文摘要 · Abstract (English)
Reward machines (RMs) provide a structured way to specify non-Markovian rewards in reinforcement learning (RL), thereby improving both expressiveness and programmability. Viewed more broadly, they separate what is known about the environment, captured by the reward mechanism, from what remains unknown and must be discovered through sampling. This separation supports techniques such as counterfactual experience generation and reward shaping, which reduce sample complexity and speed up learning. We introduce physics-informed reward machines (pRMs), a symbolic machine designed to express complex learning objectives and reward structures for RL agents, thereby enabling more programmable, expressive, and efficient learning. We present RL algorithms capable of exploiting pRMs via counterfactual experiences and reward shaping. Our experimental results show that these techniques accelerate reward acquisition during the training phases of RL. We demonstrate the expressiveness and effectiveness of pRMs through experiments in both finite and continuous physical environments, illustrating that incorporating pRMs significantly improves learning efficiency across several control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。