提出REBA框架,实现连续部分可观测环境下在线规划的可靠符号状态生成。
REBA: A Revealed Belief Automaton Framework for Online Planning in Continuous POMDPs

- 基于信息论门控动态发现噪声信念中的可靠锚点,构建符号化信念自动机
- 在巡逻与导航任务中相比顶尖方法性能提升17.0%至47.4%
- 无需预定义离散抽象,适合需形式化保障的长期逻辑规划场景
在连续部分可观测马尔可夫决策过程(POMDPs)中,基于ω-正则规范的在线规划需在有限符号内存中处理连续信念动态以追踪时序进展。现有方法或直接搜索信念空间,或依赖预设离散抽象,存在难以在噪声信念下进行可靠符号状态认证、缺乏长时序逻辑进度记忆等问题。为此,本文提出揭示信念自动机(REBA),一种事件驱动框架,将研究范式从全局信念空间离散化转向在线揭示事件的认证。具体地,提出一种在线揭示方法,通过信息论门控动态分析并建立来自连续信念空间的抽象,识别出可靠的信念锚点;进而设计增量拓扑适应机制,在经认证的锚点上构建在线有限信念自动机。结合ω-正则规范,REBA支持无需预设离散抽象的正式偶性策略合成,从而引导蒙特卡洛树搜索突破局部视野进行在线扩展。此外,设计误差分解分析以评估该离散引导对底层连续POMDP的有效性与可靠性。在巡逻与导航场景的实证评估表明,REBA在主要指标上较现有最优方法提升17.0%至47.4%。
原文摘要 · Abstract (English)
Online planning in continuous partially observable Markov decision processes (POMDPs) using $ω$-regular specifications requires handling continuous belief dynamics within the finite symbolic memory in order to track temporal progress. Existing methods based on either direct search in belief space or predefined discrete abstractions suffer from drawbacks, e.g., lack of symbolic memory for long-horizon logical progress or difficult to certify from noisy online beliefs. As such, obtaining reliable symbolic states online from continuous observations remains a challenge. To address this issue, we introduce the Revealed Belief Automaton (REBA), an event-driven framework that advances the research from global belief-space discretization to a fundamental new way of thinking, namely online certification of revelation events. Specifically, we propose an online revelation method that, through information-theoretic gates, can dynamically analyse and establish belief abstraction from the continuous belief space by discovering reliable anchors among noisy beliefs. We then develop an incremental topology adaptation mechanism over the certified anchors to realise the online finite Belief Automaton. By combining with the $ω$-regular specification, REBA is able to support formal parity policy synthesis without a predefined discrete abstraction, which in turn can guide the Monte Carlo Tree Search process to perform online search beyond its local horizon. In addition, we design an error decomposition analysis which can assess the effectiveness and reliability of this discrete guidance for the underlying continuous POMDP. Empirical evaluations in patrolling and navigation scenarios show that REBA matches or exceeds all evaluated baselines, with primary metric gains of +17.0\% to +47.4\% over state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。