arXiv:2605.01051cs.ROcs.AI2026-05被引 1

解决复杂时序逻辑任务中价值函数与策略不一致的问题。

Value Functions for Temporal Logic: Optimal Policies and Safety Filters

论文配图:Value Functions for Temporal Logic: Optimal Policies and Safety Filters
图 1 · 摘自论文原文
  • 基于状态历史构建非马尔可夫策略避免延迟完成任务
  • 证明了对嵌套Until等规范的定量鲁棒性最优
  • 将Q函数扩展为复杂时序逻辑的安全过滤器

尽管基本的到达、回避及到达-回避问题的贝尔曼方程已有深入研究,但在未折现无限时域设定下,价值最优与策略最优之间的关系变得复杂,尤其对于更复杂的任务。贪心最大化Q函数可能导致策略无限推迟任务完成,即使价值函数已是最优。基于近期将时序逻辑(TL)价值函数分解为子价值函数图的成果,我们构建了依赖状态历史的非马尔可夫策略,避免此病态现象,并证明其在嵌套Until、Globally和Globally-Until规范下的定量鲁棒性评分上最优。此外,我们进一步展示了Q函数如何作为复杂时序逻辑规范的安全过滤器,拓展了先前仅适用于简单回避或到达-回避任务的结果。

原文摘要 · Abstract (English)

While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks.

时序逻辑最优控制安全过滤强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。