arXiv:2605.12581cs.LOcs.AI2026-05中稿 · IJCAI

为不确定环境中的智能体设计可验证的逻辑导航策略

Ensuring Logic in the Fog: Sound POMDP Synthesis with LTL Objectives

论文配图:Ensuring Logic in the Fog: Sound POMDP Synthesis with LTL Objectives
图 1 · 摘自论文原文
  • 基于信念动态生成确保LTL满足的奖励信号
  • 在多个基准测试中超越现有方法,且保持可扩展性
  • 适合需要高可靠性的自主系统设计者

在部分可观测环境下合成能遵守复杂时序约束的自主智能体仍是一项根本挑战。尽管线性时序逻辑(LTL)提供了严谨的任务描述语言,但在部分可观测马尔可夫决策过程(POMDP)中定性验证LTL满足性是不可判定的,这使得量化合成困难,尤其在设计近似求解器所需的可靠奖励信号时。本文提出一种新颖的、可靠的奖励塑造机制,动态生成基于信念的奖励,其基础是经认证的LTL满足性。将该机制融入增强的蒙特卡洛规划框架后,使智能体能在‘雾’般的部分可观测环境中,通过聚焦于可验证成功的搜索过程进行导航。实验表明,该方法不仅在现有求解器失效的场景中表现优异,还在多种基准领域中保持有效性和可扩展性。

原文摘要 · Abstract (English)

Synthesising autonomous agents that can navigate uncertain environments while adhering to complex temporal constraints remains a fundamental challenge. While Linear Temporal Logic (LTL) provides a rigorous language for specifying such tasks, the inherent undecidability of qualitatively verifying LTL satisfaction in partially observable Markov decision processes renders quantitative synthesis difficult, especially when designing reliable reward signals for approximate solvers. In this paper, we bridge this gap with a novel, sound reward-shaping mechanism that dynamically generates belief-dependent rewards grounded in certified LTL satisfaction. By integrating this mechanism into an enhanced Monte Carlo Planning framework, we empower agents to navigate the `fog' of partial observability with a search process focused on maximising verifiable success. Our experiments demonstrate that this approach not only thrives in scenarios where existing solvers fail but also maintains effectiveness and scalability across diverse benchmark domains.

强化学习逻辑规划部分可观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。