将时间逻辑约束注入Transformer型强化学习,提升任务合规性。
Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies

- 用LTLf公式构建有限自动机,通过可微信号融入训练过程。
- 在导航任务中满足率提升至98.7%,同时保持与基线相当的奖励水平。
- 适用于需保证安全与可达性的高阶时序任务,如自动驾驶、机器人控制。
本文研究基于有限轨迹线性时序逻辑(LTLf)表达的长期任务约束下的离线强化学习。近期基于Transformer的方法(如轨迹变换器、决策变换器)已将强化学习建模为序列生成问题,但这些方法仅优化奖励,忽视高层时序要求。为此,本文提出一种神经符号框架,将LTLf背景知识注入此类Transformer型策略。该方法将LTLf公式编译为确定性有限自动机(DFAs),并通过可微表示和基于逻辑的损失函数将其融入学习过程。具体而言,从自动机推进中提取可微满足信号,并作为训练中的正则化项。所提方法对不同模型架构均具兼容性。在包含安全性和可达性组合性质的导航环境中评估,结果表明引入背景知识不仅显著提升约束满足率,且维持了与原始基线相当的回报性能。
原文摘要 · Abstract (English)
In this work we study offline reinforcement learning (RL) under temporally extended task constraints expressed in Linear Temporal Logic over finite traces (LTLf). Recently, transformer-based approaches such as Trajectory Transformers and Decision Transformers have been adopted to address RL as a sequence modeling problem. However, these methods optimize purely for reward and do not account for high-level temporal requirements. Here, we introduce a neurosymbolic framework that injects LTLf background knowledge into such transformer-based RL policies. Our approach compiles LTLf formulas into deterministic finite automata (DFAs) and integrates them into the learning process through a differentiable representation and a logic-based loss function. In particular, we derive differentiable satisfaction signals from DFA progression and use them as a regularization term during training. The resulting method is architecture-agnostic across different models. We evaluate the proposed framework on navigation environments with specification suites covering combinations of safety and reachability temporal properties. Experimental results show that incorporating background knowledge not only improves constraint satisfaction, but also maintains competitive return compared to vanilla baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。