arXiv:2505.18044cs.LGcs.AI2025-05NeurIPS被引 6

提出线性混合分布鲁棒强化学习框架,更精准建模动态不确定性。

Linear Mixture Distributionally Robust Markov Decision Processes

  • 用混合权重球定义不确定性集,优于传统状态动作矩形化方法
  • 在三种散度度量下证明了样本复杂度,验证可学习性
  • 适合有先验混合模型知识的鲁棒决策场景

许多现实决策问题面临离域挑战:智能体在源域学习策略后部署到状态转移不同的目标域。分布鲁棒马尔可夫决策过程(DRMDP)通过寻找在预设转移动态不确定集内最坏环境仍表现良好的鲁棒策略来应对此问题。其有效性高度依赖于基于动态先验知识设计的不确定集。本文提出一种新型线性混合分布鲁棒马尔可夫决策过程(线性混合DRMDP)框架,假设名义动态为线性混合模型。与现有直接以名义核为中心的球形不确定集不同,线性混合DRMDP基于混合权重参数的球形不确定集构建。当存在对混合模型的先验知识时,该框架相比基于(s,a)-矩形性和d-矩形性的传统模型提供了更精细的不确定性表征。本文提出了适用于一般f-散度不确定集的鲁棒策略学习元算法,并分析了在总变差、Kullback-Leibler和χ²散度三种实例下的样本复杂度。这些结果确立了线性混合DRMDP的统计可学习性,为该新设定的未来研究奠定了理论基础。

原文摘要 · Abstract (English)

Many real-world decision-making problems face the off-dynamics challenge: the agent learns a policy in a source domain and deploys it in a target domain with different state transitions. The distributionally robust Markov decision process (DRMDP) addresses this challenge by finding a robust policy that performs well under the worst-case environment within a pre-specified uncertainty set of transition dynamics. Its effectiveness heavily hinges on the proper design of these uncertainty sets, based on prior knowledge of the dynamics. In this work, we propose a novel linear mixture DRMDP framework, where the nominal dynamics is assumed to be a linear mixture model. In contrast with existing uncertainty sets directly defined as a ball centered around the nominal kernel, linear mixture DRMDPs define the uncertainty sets based on a ball around the mixture weighting parameter. We show that this new framework provides a more refined representation of uncertainties compared to conventional models based on $(s,a)$-rectangularity and $d$-rectangularity, when prior knowledge about the mixture model is present. We propose a meta algorithm for robust policy learning in linear mixture DRMDPs with general $f$-divergence defined uncertainty sets, and analyze its sample complexities under three divergence metrics instantiations: total variation, Kullback-Leibler, and $χ^2$ divergences. These results establish the statistical learnability of linear mixture DRMDPs, laying the theoretical foundation for future research on this new setting.

强化学习鲁棒决策分布鲁棒线性混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。