解决人机系统中多动态模型不确定性下的决策难题
Resolving Multiple-Dynamic Model Uncertainty in Hypothesis-Driven Belief-MDPs
- 构建可同时推理多个假设的信念马尔可夫决策过程
- 在确定正确模型与优化控制表现间取得平衡
- 支持稀疏树搜索,适用于连续状态空间
当人机系统操作员遇到意外行为时,常需考虑多个可能解释。通过信息采集动作(如额外测量或控制输入)可缓解不确定性并确定最准确假设。该任务可建模为假设驱动的信念马尔可夫决策过程(hypothesis-driven belief MDP)。然而,此问题面临类似部分可观测马尔可夫决策过程(POMDP)的“历史诅咒”:在连续域中,需处理无穷多可能的动作-观测历史,每条历史对应一个不同的状态信念。在假设驱动场景下,每个动作-观测对会为每个假设生成不同信念,导致额外分支。本文针对每个假设对应一个不同动态模型的情形,提出一种新信念MDP形式化方法:(i) 支持多假设推理;(ii) 平衡确定最可能正确假设与在底层POMDP中表现良好;(iii) 可通过稀疏树搜索求解。
原文摘要 · Abstract (English)
When human operators of cyber-physical systems encounter surprising behavior, they often consider multiple hypotheses that might explain it. In some cases, taking information-gathering actions such as additional measurements or control inputs given to the system can help resolve uncertainty and determine the most accurate hypothesis. The task of optimizing these actions can be formulated as a belief-space Markov decision process that we call a hypothesis-driven belief MDP. Unfortunately, this problem suffers from the curse of history similar to a partially observable Markov decision process (POMDP). To plan in continuous domains, an agent needs to reason over countlessly many possible action-observation histories, each resulting in a different belief over the unknown state. The problem is exacerbated in the hypothesis-driven context because each action-observation pair spawns a different belief for each hypothesis, leading to additional branching. This paper considers the case in which each hypothesis corresponds to a different dynamic model in an underlying POMDP. We present a new belief MDP formulation that: (i) enables reasoning over multiple hypotheses, (ii) balances the goals of determining the (most likely) correct hypothesis and performing well in the underlying POMDP, and (iii) can be solved with sparse tree search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。