arXiv:2511.12808cs.LGcs.AI2025-11AAAI

用时序逻辑生成密集奖励,让智能体更快学会复杂任务。

Expressive Temporal Specifications for Reward Monitoring

  • 用定量时序逻辑构造可实时计算的奖励监控器
  • 在多个环境中减少收敛时间并提升任务完成度
  • 适合长周期决策、奖励稀疏的任务场景

在强化学习中,设计有效且密集的奖励函数仍是关键挑战,直接影响训练效率。本文利用有限轨迹上的定量线性时序逻辑($ ext{LTL}_f[ ext{F}]$)的表达能力,合成能够对运行时可观测状态轨迹生成密集奖励流的奖励监控器。通过提供细粒度反馈,这些监控器引导智能体趋向最优行为,缓解了长周期决策下因布尔语义导致的奖励稀疏问题。该框架与算法无关,仅依赖状态标记函数,天然支持非马尔可夫性质的建模。实验表明,定量监控器在最大化任务完成度量化指标和减少收敛时间方面,始终优于甚至超越布尔监控器。

原文摘要 · Abstract (English)

Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we harness the expressive power of quantitative Linear Temporal Logic on finite traces (($\text{LTL}_f[\mathcal{F}]$)) to synthesize reward monitors that generate a dense stream of rewards for runtime-observable state trajectories. By providing nuanced feedback during training, these monitors guide agents toward optimal behaviour and help mitigate the well-known issue of sparse rewards under long-horizon decision making, which arises under the Boolean semantics dominating the current literature. Our framework is algorithm-agnostic and only relies on a state labelling function, and naturally accommodates specifying non-Markovian properties. Empirical results show that our quantitative monitors consistently subsume and, depending on the environment, outperform Boolean monitors in maximizing a quantitative measure of task completion and in reducing convergence time.

强化学习奖励设计时序逻辑稠密奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。