arXiv:2507.02942eess.SYcs.AI2025-07中稿 · publication in the…被引 1

为部分可观测环境中的复杂感知任务设计最优控制策略

Control Synthesis in Partially Observable Environments for Complex Perception-Related Objectives

  • 用线性不等式时序逻辑定义感知目标,结合信念空间约束
  • 将复杂目标转为可达性问题,通过信念马尔可夫决策过程求解
  • 结合蒙特卡洛树搜索提升可扩展性,适合无人机探测等场景

自主系统在部分可观测环境下常面临感知相关任务。本文研究在部分可观测马尔可夫决策过程(POMDP)建模的环境中,针对复杂感知目标合成最优策略的问题。为形式化表达此类目标,提出一种新的时序逻辑——共安全线性不等式时序逻辑(sc-iLTL),可将原子命题的逻辑组合表示为对信念空间的线性不等式约束。通过构建sc-iLTL目标对应的确定有限自动机,并与信念MDP进行乘积构造,将sc-iLTL目标转化为可达性目标。为应对乘积结构带来的可扩展性挑战,引入一种收敛于最优策略的概率性蒙特卡洛树搜索(MCTS)方法。最后,通过无人机探测案例验证了该方法的有效性。

原文摘要 · Abstract (English)

Perception-related tasks often arise in autonomous systems operating under partial observability. This work studies the problem of synthesizing optimal policies for complex perception-related objectives in environments modeled by partially observable Markov decision processes. To formally specify such objectives, we introduce \emph{co-safe linear inequality temporal logic} (sc-iLTL), which can define complex tasks that are formed by the logical concatenation of atomic propositions as linear inequalities on the belief space of the POMDPs. Our solution to the control synthesis problem is to transform the \mbox{sc-iLTL} objectives into reachability objectives by constructing the product of the belief MDP and a deterministic finite automaton built from the sc-iLTL objective. To overcome the scalability challenge due to the product, we introduce a Monte Carlo Tree Search (MCTS) method that converges in probability to the optimal policy. Finally, a drone-probing case study demonstrates the applicability of our method.

强化学习部分可观测形式化验证无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。