arXiv:2510.23744cs.AI2025-10NeurIPS被引 2

解决多个不确定环境下的鲁棒决策问题,找到对最坏情况也有效的策略。

Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability

  • 将多环境模型扩展为初始信念集,构建对抗性信念POMDP框架。
  • 证明任意多环境模型可简化为仅变转移或观测函数的等效形式。
  • 提出精确与近似算法,适用于标准基准测试的多环境强化学习任务。

多环境部分可观测马尔可夫决策过程(ME-POMDP)在标准POMDP基础上引入离散模型不确定性。它表示一组共享状态、动作和观测空间但转移、观测和奖励模型可任意不同的POMDP集合。这类模型常见于多个领域专家对问题建模存在分歧时。目标是寻找一个能在所有可能的POMDP中表现最优的单一鲁棒策略,即最大化最坏情况下的累积奖励。本文进一步拓展已有工作:首先,将ME-POMDP推广至初始信念集情形,称为对抗性信念POMDP(AB-POMDP);其次,证明任意ME-POMDP可等价转化为仅转移与奖励或仅观测与奖励变化的简化形式,且保持最优策略不变。随后,我们设计了用于求解AB-POMDP的精确与近似(基于点)算法,从而支持对标准POMDP基准问题在多环境设定下的鲁棒策略计算。

原文摘要 · Abstract (English)

Multi-environment POMDPs (ME-POMDPs) extend standard POMDPs with discrete model uncertainty. ME-POMDPs represent a finite set of POMDPs that share the same state, action, and observation spaces, but may arbitrarily vary in their transition, observation, and reward models. Such models arise, for instance, when multiple domain experts disagree on how to model a problem. The goal is to find a single policy that is robust against any choice of POMDP within the set, i.e., a policy that maximizes the worst-case reward across all POMDPs. We generalize and expand on existing work in the following way. First, we show that ME-POMDPs can be generalized to POMDPs with sets of initial beliefs, which we call adversarial-belief POMDPs (AB-POMDPs). Second, we show that any arbitrary ME-POMDP can be reduced to a ME-POMDP that only varies in its transition and reward functions or only in its observation and reward functions, while preserving (optimal) policies. We then devise exact and approximate (point-based) algorithms to compute robust policies for AB-POMDPs, and thus ME-POMDPs. We demonstrate that we can compute policies for standard POMDP benchmarks extended to the multi-environment setting.

强化学习部分可观测鲁棒控制POMDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。