用采样与模型检测结合,为不确定环境决策生成有保证的控制器。
Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning

- 以采样作查询,模型检测作验证,学习有限状态控制器
- 在采样策略为正则时,可证明相对完备性并解决安全阈值问题
- 适合需要形式化保证但传统方法难处理的复杂决策场景
部分可观测马尔可夫决策过程(POMDP)是不确定性环境下决策的标准框架。采样方法虽能良好扩展,但缺乏形式正确性保证,不适用于安全关键任务;而形式化合成技术虽能保证正确性,但普遍面临可扩展性问题,因一般POMDP合成是不可判定的。为弥合这一差距,我们提出一种融合采样、自动机学习与模型检测的合成框架。受Angluin的$L^*$算法启发,该方法将采样视为成员查询,模型检测作为等价查询,从而在采样诱导策略为正则的前提下,合成具有形式化保证的有限状态控制器。我们建立了该框架的相对完备性结果。原型实现的实验表明,该方法成功解决了现有形式化合成工具仍难以处理的阈值安全问题。我们认为该算法可作为应对POMDP合成固有困难的组合式方法中的重要组件。
原文摘要 · Abstract (English)
Partially Observable Markov Decision Processes (POMDPs) are the standard framework for decision-making under uncertainty. While sampling-based methods scale well, they lack formal correctness guarantees, making them unsuitable for safety-critical applications. Conversely, formal synthesis techniques provide correctness-by-construction but often struggle with scalability, as general POMDP synthesis is undecidable. To bridge this gap, we propose a synthesis framework that integrates sampling, automata learning, and model-checking. Inspired by Angluin's $L^*$ algorithm, our approach utilizes sampling as a membership oracle and model-checking as an equivalence oracle. This enables the synthesis of finite-state controllers with formal guarantees, provided the sampling-induced policy is regular. We establish a relative completeness result for this framework. Experimental results from our prototypical implementation demonstrate that this method successfully solves threshold-safety problems that remain challenging for existing formal synthesis tools. We believe our algorithm serves as a valuable component in a portfolio approach to tackling the inherent difficulty of POMDP synthesis problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。