用辅助信息提升迁移学习在环境变化下的鲁棒性
Robust Transfer Learning with Side Information
- 基于有限目标样本与源-目标动态先验构建估计中心的不确定性集
- 在多个环境上优于现有鲁棒与非鲁棒基线,提升策略性能
- 适用于低维结构转移场景,显著减少策略次优差距
鲁棒马尔可夫决策过程通过分布鲁棒优化(DRO)在转移核的不确定集内寻找最差情况下的最优策略来应对环境变化。然而,标准DRO方法在大环境变化下需扩大不确定性集,导致策略过于保守。本文提出一种迁移学习框架,通过结合少量目标样本与关于源-目标动态的辅助信息(如特征矩、分布距离、密度比的边界),构建估计中心的不确定性集,实现更精准的核估计与更紧的不确定性集。建立了鲁棒与非鲁棒值函数的误差界与收敛性结果,并提供了学习到的鲁棒策略的有限样本保证及鲁棒次优性差距分析。在转移模型具有轻微低维结构时,辅助信息能缩小该差距并提升样本效率。在OpenAI Gym与经典控制任务上的实验表明,该方法在目标域表现持续优于最先进的鲁棒与非鲁棒基线。
原文摘要 · Abstract (English)
Robust Markov Decision Processes (MDPs) address environmental shift through distributionally robust optimization (DRO) by finding an optimal worst-case policy within an uncertainty set of transition kernels. However, standard DRO approaches require enlarging the uncertainty set under large shifts, which leads to overly conservative and pessimistic policies. In this paper, we propose a framework for transfer under environment shift that derives a robust target-domain policy via estimate-centered uncertainty sets, constructed through constrained estimation that integrates limited target samples with side information about the source-target dynamics. The side information includes bounds on feature moments, distributional distances, and density ratios, yielding improved kernel estimates and tighter uncertainty sets. The side information includes bounds on feature moments, distributional distances, and density ratios, yielding improved kernel estimates and tighter uncertainty sets. Error bounds and convergence results are established for both robust and non-robust value functions. Moreover, we provide a finite-sample guarantee on the learned robust policy and analyze the robust sub-optimality gap. Under mild low-dimensional structure on the transition model, the side information reduces this gap and improves sample efficiency. We assess the performance of our approach across OpenAI Gym environments and classic control problems, consistently demonstrating superior target-domain performance over state-of-the-art robust and non-robust baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。