提出一种新启发式,显著加速不确定环境下的任务规划。
Heuristics for Partially Observable Stochastic Contingent Planning
- 基于领域结构构建启发式,融合信息价值与随机性建模
- 计算开销略增,但所需轨迹数减少一个数量级
- 适合需要大量信息收集的复杂决策场景
在随机且部分可观测的环境中完成任务是人工智能中的重要问题,通常被建模为基于目标的POMDP。可通过RTDP-BEL算法求解,该算法通过从初始信念到目标的前向轨迹运行来推进。轨迹可由启发式引导,更准确的启发式能显著加快收敛速度。本文提出一种利用领域模型结构表示的启发式函数:在松弛空间中计算达成目标的计划,同时考虑信息价值和随机效应。实验表明,尽管该启发式计算较慢,但收敛所需轨迹数减少一个数量级,从而整体提升RTDP-BEL效率,尤其适用于需大量信息获取的问题。
原文摘要 · Abstract (English)
Acting to complete tasks in stochastic partially observable domains is an important problem in artificial intelligence, and is often formulated as a goal-based POMDP. Goal-based POMDPs can be solved using the RTDP-BEL algorithm, that operates by running forward trajectories from the initial belief to the goal. These trajectories can be guided by a heuristic, and more accurate heuristics can result in significantly faster convergence. In this paper, we develop a heuristic function that leverages the structured representation of domain models. We compute, in a relaxed space, a plan to achieve the goal, while taking into account the value of information, as well as the stochastic effects. We provide experiments showing that while our heuristic is slower to compute, it requires an order of magnitude less trajectories before convergence. Overall, it thus speeds up RTDP-BEL, particularly in problems where significant information gathering is needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。