arXiv:2502.18296cs.GTcs.AI2025-02被引 2

研究多目标强化学习中如何用有限策略组合逼近任意收益,揭示随机策略的必要性。

Mixing Any Cocktail with Limited Ingredients: On the Structure of Payoff Sets in Multi-Objective POMDPs and its Impact on Randomised Strategies

  • 通过混合有限个纯策略,可逼近任意期望收益向量。
  • 对所有策略下期望收益有界的场景,可精确实现任意期望收益。
  • 为多目标部分可观测决策问题提供随机策略设计依据,适合强化学习研究者。

我们研究部分可观测马尔可夫决策过程(POMDPs)中的多维收益函数。分析所有策略(政策)所能产生的期望收益向量的结构,并探讨达到特定期望收益向量所需策略类型。一般情况下,仅使用纯策略(不依赖随机化)不足以实现目标。我们证明:对于所有策略下期望定义良好的收益,只需混合有限个纯策略,即可在任意精度内逼近任一期望收益向量。此外,对于所有策略下期望收益有限的收益,任意期望收益均可通过混合有限个策略精确实现。

原文摘要 · Abstract (English)

We consider multi-dimensional payoff functions in partially observable Markov decision processes. We study the structure of the set of expected payoff vectors of all strategies (policies) and study what kind are needed to achieve a given expected payoff vector. In general, pure strategies (i.e., not resorting to randomisation) do not suffice for this problem. We prove that for any payoff for which the expectation is well-defined under all strategies, it is sufficient to mix (i.e., randomly select a pure strategy at the start of a play and committing to it for the rest of the play) finitely many pure strategies to approximate any expected payoff vector up to any precision. Furthermore, for any payoff for which the expected payoff is finite under all strategies, any expected payoff can be obtained exactly by mixing finitely many strategies.

多目标强化学习部分可观测策略混合期望收益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。