针对有限时域非平稳强化学习,提出可迁移的深度Q-learning方法。
Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning
- 用重加权策略构造可迁移的强化学习样本
- 在小样本非平稳场景下提升决策性能
- 适合医疗、商业等动态决策领域应用
在商业和医疗等动态决策场景中,利用来自不同人群的样本轨迹可显著提升特定目标人群的强化学习性能,尤其在样本量有限时。现有迁移学习方法多聚焦于线性回归,难以直接应用于强化学习。本文首次研究基于神经网络的非平稳有限时域马尔可夫决策过程中的迁移学习,采用后向归纳学习。我们发现,在回归中有效的简单样本合并策略在马尔可夫决策过程中失效。为此,提出「重加权靶向过程」以构建「可迁移的RL样本」,并引入「迁移深度Q*学习」,实现带理论保障的神经网络近似。假设奖励函数可迁移,处理转移密度可迁移或不可迁移两种情形。所提分析技术对神经网络的迁移学习及领域偏移场景具更广意义。合成与真实数据集上的实验验证了方法优势,展示了通过有策略构建可迁移样本改善非平稳强化学习决策的潜力。
原文摘要 · Abstract (English)
In dynamic decision-making scenarios across business and healthcare, leveraging sample trajectories from diverse populations can significantly enhance reinforcement learning (RL) performance for specific target populations, especially when sample sizes are limited. While existing transfer learning methods primarily focus on linear regression settings, they lack direct applicability to reinforcement learning algorithms. This paper pioneers the study of transfer learning for dynamic decision scenarios modeled by non-stationary finite-horizon Markov decision processes, utilizing neural networks as powerful function approximators and backward inductive learning. We demonstrate that naive sample pooling strategies, effective in regression settings, fail in Markov decision processes.To address this challenge, we introduce a novel ``re-weighted targeting procedure'' to construct ``transferable RL samples'' and propose ``transfer deep $Q^*$-learning'', enabling neural network approximation with theoretical guarantees. We assume that the reward functions are transferable and deal with both situations in which the transition densities are transferable or nontransferable. Our analytical techniques for transfer learning in neural network approximation and transition density transfers have broader implications, extending to supervised transfer learning with neural networks and domain shift scenarios. Empirical experiments on both synthetic and real datasets corroborate the advantages of our method, showcasing its potential for improving decision-making through strategically constructing transferable RL samples in non-stationary reinforcement learning contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。