arXiv:2508.00707cs.LGcs.AI2025-08AAAI被引 1

通过因子化结构提升鲁棒MDP的学习效率和性能

Efficient Solution and Learning of Robust Factored MDPs

  • 利用系统组件间的独立性建模,将复杂优化转为可解线性规划
  • 实验表明样本效率显著提升,政策性能更优且保证更紧
  • 适合需要高可靠性的强化学习场景,如机器人控制

鲁棒马尔可夫决策过程(r-MDP)通过显式建模对转移动态的认知不确定性来扩展标准MDP。从与未知环境的交互中学习r-MDP可合成具有可证明(PAC)性能保障的鲁棒策略,但通常需大量样本。本文提出基于因子化状态空间表示的新方法,利用系统各组件间模型不确定性的独立性。尽管因子化r-MDP的策略合成导致难以求解的非凸优化问题,我们展示了如何将其重构成可处理的线性规划。在此基础上,还提出了直接学习因子化模型表示的方法。实验结果表明,利用因子化结构可实现维度上的样本效率提升,生成比现有最优方法更有效的鲁棒策略,并具备更紧的性能保证。

原文摘要 · Abstract (English)

Robust Markov decision processes (r-MDPs) extend MDPs by explicitly modelling epistemic uncertainty about transition dynamics. Learning r-MDPs from interactions with an unknown environment enables the synthesis of robust policies with provable (PAC) guarantees on performance, but this can require a large number of sample interactions. We propose novel methods for solving and learning r-MDPs based on factored state-space representations that leverage the independence between model uncertainty across system components. Although policy synthesis for factored r-MDPs leads to hard, non-convex optimisation problems, we show how to reformulate these into tractable linear programs. Building on these, we also propose methods to learn factored model representations directly. Our experimental results show that exploiting factored structure can yield dimensional gains in sample efficiency, producing more effective robust policies with tighter performance guarantees than state-of-the-art methods.

强化学习鲁棒控制因子化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。