arXiv:2602.06227cs.AIcs.LG2026-02AAAI被引 2

用一阶时序逻辑直接定义复杂奖励,让强化学习更懂人类意图。

Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement Learning (Extended Version)

  • 基于一阶时序逻辑构建可表达复杂任务的奖励机制
  • 在连续控制任务中实现高成功率的复杂目标求解
  • 适合需要精确语义描述的任务,如机器人操作

本文提出一种新框架,用于在大规模状态空间的马尔可夫决策过程(MDPs)中对非马尔可夫奖励进行逻辑建模。方法采用有限迹线上的线性时序逻辑模理论(LTLfMT),将谓词扩展为任意一阶理论的公式,而非简单布尔变量,从而支持对非结构化异构数据域中复杂任务的自然描述。该增强表达力虽带来理论与计算挑战,但本文从理论上识别出一个在无限状态空间下既可计算又足够表达的LTLfMT子类。实践中,结合奖励机器与后见经验回放(HER)方法,实现了对一阶逻辑规范的高效转化,并缓解奖励稀疏问题。实验基于非线性算术理论,在连续控制场景中验证了该方法的有效性,结果表明定制化的HER实现对解决复杂目标至关重要。

原文摘要 · Abstract (English)

In this work, we propose a novel framework for the logical specification of non-Markovian rewards in Markov Decision Processes (MDPs) with large state spaces. Our approach leverages Linear Temporal Logic Modulo Theories over finite traces (LTLfMT), a more expressive extension of classical temporal logic in which predicates are first-order formulas of arbitrary first-order theories rather than simple Boolean variables. This enhanced expressiveness enables the specification of complex tasks over unstructured and heterogeneous data domains, promoting a unified and reusable framework that eliminates the need for manual predicate encoding. However, the increased expressive power of LTLfMT introduces additional theoretical and computational challenges compared to standard LTLf specifications. We address these challenges from a theoretical standpoint, identifying a fragment of LTLfMT that is tractable but sufficiently expressive for reward specification in an infinite-state-space context. From a practical perspective, we introduce a method based on reward machines and Hindsight Experience Replay (HER) to translate first-order logic specifications and address reward sparsity. We evaluate this approach to a continuous-control setting using Non-Linear Arithmetic Theory, showing that it enables natural specification of complex tasks. Experimental results show how a tailored implementation of HER is fundamental in solving tasks with complex goals.

强化学习时序逻辑奖励设计机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。